Nobody opens a spreadsheet mid shift. They send a message, misspelled, with the short code, two bottles in one sentence. The console below is the real resolver, running in your browser. Type into it.
Type the way someone types at 1am with a rag in the other hand. Misspell it, use the short code, put two bottles in one sentence. The trace under each line is the real resolution path, running in your browser.
jd6 is six Jack Daniels
sold 2 tito is a misspelling worth 0.89
sold 40 aperol is more than the bar has
A language model classifies what the bartender is trying to do and stops there. Which bottle and how many are decided by code that returns the same answer every time. Every version where we let the model make that call was confidently wrong often enough to be useless, and a bar cannot audit a hallucinated count at close.
Resolution runs in order and gets more forgiving as it goes. Normalize strips accents, case and punctuation. Synonym catches the spellings people actually use. Exact handles the clean case. Only then does fuzzy score the input against every brand on the list. Free text gets one more pass, a one to four word window sliding across the message so Don Julio and Casa Amigo Reposado survive being buried mid sentence.
No single signal can move inventory on its own. That threshold is the boring half of the design. The interesting half is the boundary above it: pour more than the bar holds and the write is held rather than applied.
That is not error handling. It is the product decision the whole system is built around, which is what gets applied quietly and what a person has to look at. Try sold 40 aperol in the console and watch it refuse.
A bottle on its own is 0.4 and moves nothing. A bottle with a number is 0.7 and lands exactly on the line.
The cascade shipped without an evaluation harness. We watched it work in the bar and tuned the fuzzy floor by feel, which is enough to launch and not enough to improve, because you cannot tell a better threshold from a worse one without a fixed set of inputs and known answers to score against. So we built one: 133 labelled bartender messages covering misspellings, short codes, two bottles in one sentence, near collisions like Don Julio against Don Julio Reposado, and messages that must move nothing at all.
It found four defects on the first run, and the score went from 90.2% to 97.7% once they were fixed. Overlapping windows were counting one pour as two bottles, because Don Julio Reposado also matches Don Julio. Catalogue names of three or four letters were matching ordinary English, so the word do resolved to Dom at exactly the floor. The short code expander was eating part of longer names and turning Mr. Black into Martini Rosso. And a forecast question phrased as what needs restocking was falling through to small talk.
97.7%
still failing
| floor | recall | non bottles let in |
|---|---|---|
| 0.70 | 82.2% | 1 |
| 0.75 | 81.3% | 0 |
| 0.80 shipped | 80.4% | 0 |
| 0.85 | 79.4% | 0 |
| 0.90 | 73.8% | 0 |
Loosening the floor buys a little recall and immediately starts letting things that are not bottles through. Tightening it costs recall for nothing. The number we picked by feel turned out to be roughly where it belongs, which we only know now.
It also killed our favourite theory. We had assumed short brand names were the weak spot and that the threshold should vary by name length. Recall by name length came back flat, so that was wrong. Name length does matter, but on the other side of the ledger: short names let ordinary words in rather than keeping real ones out, which is a different fix and one the guess would have missed.
Three cases still fail and we left them failing. Two are judgement calls about how far a surname or a digit swapped for a word should be trusted, and one is bar slang where 86 means we are out rather than eighty six bottles. The harness runs on every change with a floor under the score, so the next person to touch this finds out immediately rather than in a bar at closing time.
Bottlely was two founders and three part time engineers, and this page is the part we built: the resolution cascade, the boundary that keeps the model out of the write path, the chat client and the Flask service behind it.
We stopped selling to bars. Takunda took the engine into a new company called Ventory, pointed at e-commerce inventory, and runs it now. The piece worth carrying into a different industry turned out to be this one, the layer that turns a mess of human input into something a business can act on.