Eligibility engine for point-of-sale BNPL
Digital Payment Systems — leading developer
Leading development of the customer eligibility decision engine for a point-of-sale buy-now-pay-later product, in development ahead of production launch.
In development, ahead of production launch. I lead the eligibility engine: the component that decides whether a Chase customer is even eligible for a point-of-sale loan before an offer is ever shown.
The interesting constraint is the one that makes it hard. Every decision has a 200ms p99 latency budget, and every decision must remain explainable long after it was made.
The decision path
A determination is the intersection of two inputs:
- Risk signals, arriving as records on Kafka from in-house risk services.
- Product rules, authored by the product team and versioned over time.
The engine consumes the risk records, evaluates them against the current rule set, and returns a decision. The budget covers the whole path: consume, evaluate, decide, record. There is no room in that 200ms for a synchronous fallback to a service that might be slow, because a fallback that sometimes exceeds the budget is worse than a fast failure.
Holding the budget
Meeting 200ms reliably is mostly a caching problem, because the expensive work is the same work every time.
- Eligibility decisions from risk are cached. Risk signals for a given customer do not change between two offers seconds apart, so recomputing them is wasted work.
- Offer plans from pricing are cached for the same reason. Pricing is deterministic given an input set, so a plan computed once stays valid.
The subtlety is invalidation. A cache that is never invalidated returns stale answers, and a stale eligibility answer is a compliance problem rather than a performance one. So cache entries carry the rule version that produced them, and a rule change invalidates by version rather than by wall-clock expiry. The cache can be wrong for at most as long as a rule version remains current, and that window is explicit.
Making a decision reproducible
This is the part I would defend hardest in an interview.
When a customer is declined, the system stores the step that caused the failure and the reason code for it. When a customer is approved, it stores every offer plan that was generated, with the selected one marked. These are separate audit tables, and they answer a question that is much harder than it sounds: six months from now, why did this customer receive this offer?
By then the rules have changed, the pricing has changed, and the risk services have changed. You cannot reconstruct the decision by re-running the engine. You can only answer it if you recorded, at decision time, what the inputs were and which rules applied to them.
That is the whole design constraint. A decision engine that is cheap to run and impossible to audit is a liability, because the moment there is a dispute someone has to reverse-engineer your system under time pressure. Recording the failing step and the generated plans at decision time is cheap — it is a write on a path that is already writing — and it converts an archaeology exercise into a lookup.
Why this is durable-execution work
Strip away the banking vocabulary and the problem is familiar: take events from a log, make a decision from them under a latency budget, and keep enough state to explain what happened. The decision must be reproducible even though the system producing it has changed underneath. That is the same durability guarantee a workflow engine provides, applied to a decision rather than a workflow, and it is why I am targeting this kind of team.
The launch is next summer. Until then the honest description of this work is that it is in development, and I would rather say that than imply otherwise.