
Why our refund flow never asks the model for permission
Language models are good at conversation and bad at policy. Here is how we split the two, and what it did to our error rate.
By Tomasz Wrona
Engineering notes, research and playbooks from the team building Synth.

Language models are good at conversation and bad at policy. Here is how we split the two, and what it did to our error rate.
By Tomasz Wrona

Reviewing fifty tickets a week felt rigorous. Scoring all of them showed us how much the sample was hiding.

People notice a pause on the phone at around a second. This is the latency budget we work to, stage by stage.

The plan we run with every new customer, week by week, including what we check before real traffic arrives.

Most quality rubrics are written for people who already know what good looks like. Here is how to write one for a model.

One model per agent is simple and expensive. One model per step is cheaper, faster and easier to change. Here is how routing works.

An agent that never hands over is either doing very easy work or doing hard work badly. How we design the moment it passes to a person.

Synthetic test cases miss the strange things customers actually say. A step-by-step guide to turning your history into tests.

When a winter storm grounded two hundred flights, an airline's agent handled the morning. What we saw from the inside.