Skip to content
View abd-ulbasit's full-sized avatar

Block or report abd-ulbasit

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
abd-ulbasit/README.md

Abdul Basit Sajid

Backend and infrastructure engineer, working remotely with a UK team since 2023. I work on capacity, latency and the deployment path: what a system costs to run, why it is slow, and how it ships without going down. Go, Python, Kubernetes, Postgres.

Four times the number was not what it looked like.

A copy-on-write system copied what it read

pgoverlay gives each pull request its own Postgres branch: a stock container whose data directory is an OverlayFS view of one shared seed, meant to store only what that branch changes. A branch of a 5 GiB database held 5.05 GiB before its first query, and one SELECT count(*) copied a 488.5 MiB table into a branch that had only read it.

Same cause both times. Postgres opens files O_RDWR even when it only reads them: every file in the data directory before WAL replay, because not every platform allows fsync() on a read-only descriptor, and every table file a query touches. OverlayFS copies a file up on the open, not on the write. A one-flag control run proved the first (recovery_init_sync_method=syncfs: 16 KiB, not 5.05 GiB). For the second, v1.0.0 preloads a small library that opens table files read-only until their first write, so reads copy no table data, at a warm pgbench cost of at most 1.6% at the median.

The 5 GiB numbers come from a Colima VM on an M1 Pro (create times are the median of five runs), the 488.5 MiB read from Docker on Linux with ext4, and the pgbench numbers from GitHub-hosted amd64 and arm64 runners. All of them, with their raw runs, are in the benchmarks doc, and the write-up tells the whole story.

The hazard was documented one layer below where it bit me

In goqueue, Log.ReadFrom took a read lock and then called an exported method that took the same lock again. Go's RWMutex is not reentrant, and a pending writer blocks new readers, so the second acquisition deadlocks against a writer that is itself waiting on the first.

I had written the warning myself, one layer down:

// Note: We need to unlock before calling ReadFrom (it locks again)

Knowing a hazard and finding every instance of it are separate jobs. There was a second one in the quota manager.

A metric that reads as broken and is correct

sluice exports sluice_pool_hits_total, and on the L4 path it is always zero. Making it move means not forwarding the client's FIN to the backend. I tried that. A client that half-closed received an empty response, and because the abandoned socket went back into the pool still owing a reply, the next client could be served the previous client's response body.

Both were reproduced, the change was reverted, and the README explains the zero instead. An idle counter is cheaper than a protocol-visible correctness bug.

Eight milliseconds on the GPU, 2,400 Python calls around it

A real-time inference service took 2.8 s per request while GPU utilisation sat between 0 and 13% under normal load, reaching 83% only at 150 concurrent users. Inference itself was 8 ms per model, and around each of those 8 ms were roughly 2,400 Python function calls. A bigger GPU makes the 8 ms smaller and does nothing at all to the rest, which is why buying hardware had not helped: the ceiling was a single Python process, not the accelerator. Fixing the fan-out is what made four GPUs usable in the first place. Proved to 500 concurrent under load test, on four L40Ses.


Some of this is still wrong. I do not know which part yet.

basit.engineer · pgoverlay · goqueue · steward · sluice · forgepoint · Kubernetes guide

Pinned Loading

  1. steward steward Public

    Kubernetes operator (kubebuilder v4). A 20-line Application CRD provisions Postgres, Redis and RabbitMQ, injects credentials into the workload, and reports configuration drift by comparing spec gen…

    Go

  2. forgepoint forgepoint Public

    ML-lifecycle platform: 11 Go microservices on Kubernetes over gRPC and NATS JetStream, each implementing one distributed-systems pattern to production depth — saga compensation, CQRS, event sourcin…

    Go

  3. sluice sluice Public

    L4/L7 proxy and API gateway in Go. Weighted round-robin and least-connections balancing, token-bucket rate limiting, and a three-state circuit breaker with active health checks. No dependencies in …

    Go 1

  4. goqueue goqueue Public

    Distributed message queue combining Kafka's durability + SQS's simplicity + features neither has

    Go 2

  5. bookstore-kubernetes-guide bookstore-kubernetes-guide Public

    A standalone, hands-on Kubernetes guide that takes you from zero to production by progressively building, deploying, scaling, securing, observing, and operating one realistic microservices applicat…

    HCL