Mekhi Thomas-El
Software Engineer, Distributed Systems
What I build
Systems that stay correct under failure, and that can prove they did. Most of my work sits in two places: the low-level storage and filesystem layer, where correctness is enforced by the code itself, and the distributed decisioning layer, where a request has a hard latency budget and a decision that must be reproducible months later.
Those are the same problem at different scales. A filesystem that can migrate an on-disk format without going down is solving an exactly-once problem. A payment eligibility decision that must be re-explainable a year later is solving the same one again.
Selected work
Kafka store-and-forward broker
Store-and-forward message broker that took delivery reliability from 98.2% to 99.95% and cut incident-driven outages by 14 per quarter.
Eligibility engine for point-of-sale BNPL
Leading development of the customer eligibility decision engine for a point-of-sale buy-now-pay-later product, in development ahead of production launch.
Multi-part file cache rewrite
Rewrote the operating system filesystem caching layer to support multi-part files, improving data retrieval performance across storage operations.
File scanner for multi-part splitting
Implemented the scanner that decides which files are worth splitting into parts, and tuned its default behaviour for scan time and efficiency across the installed fleet.
Things I build
Side projects, mostly for the exercise of picking up an unfamiliar stack properly. Each one is a different answer to a different question.
A job search treated as a pipeline
Multi-tenant by design from the first migration, so no table is ever one missing scope filter away from leaking one user's data into another's. Résumés are versioned per application, because the question "which version did I send them" has to be answerable a month later. Six architecture decision records, including why ATS ingestion is best-effort with a manual fallback rather than the only route in.
Local first, remote enrichment, nothing overwritten
A music tagger that never sends your library anywhere. The interesting part is restraint: per-field confidence scoring so a matched artist is not trusted like a guessed year, and a review queue instead of automatic writes, because auto-tagging a collection you care about is a good way to ruin it. A three-crate workspace with the core independent of both front ends, and nine ADRs.
Rules where the obvious implementation is wrong
A chess rules engine behind a protobuf service, written to get properly into value semantics and exhaustive testing. The engine has to distinguish a piece that can reach a square from one that may move there, because moving a pinned piece exposes the king — a distinction most implementations miss and casual testing never catches. 143 tests across the engine.
Prove it, don't read it
Retrying a request must not produce a second decision. That is the property the eligibility engine depends on, and it is easy to claim and easy to get wrong, so this one is interactive.
One click sends two requests under the same idempotency key. The second carries a different customer and amount, so a matching response cannot be explained by the inputs being the same.
Rate limited, and the store holds at most 1024 keys in memory, so it resets when the server restarts.
This server, right now
Every number below is scraped from this process. No mock data, no screenshots.
Raw text exposition: /metrics · JSON: /api/stats
How this is built
A single Go binary using nothing but the standard library. No web framework, no client library, no markdown parser, no charting library. The Prometheus exposition format, the histogram quantiles, the markdown renderer, and the chart are all implemented here, because depending on them would defeat the point of the page.
net/httprouting and handlers, no framework- Hand-written Prometheus text exposition, validated against the reference parser
- Lock-sharded counters and histograms, safe under
go test -race - Idempotency-key store for the decision endpoint, in memory