Project breakdown · 2026
Guess the day’s dish.
Jacob Poteet Senior Technical Narrative Designer. Solo creator of Lunch Special.
Every guess is a real dish, and the kitchen tells you what that dish shares with the Special. I designed and implemented the gameplay loop, the CMS for the menu of 364 dishes, the telemetry underneath, and the experiment log.
Serving Special No. — · next one in —
Read live from the production database when this page loads.
The short version
A dish has no letter slots, so Wordle’s colours mean nothing here. I designed a feedback system for food and ran two channels at once, one for the player who knows the dish and one for the player who has never heard of it.
Design → 02One beat sheet specifies every clue. The same five beats, 1,820 clues deep, so a dish a day stays a defined job instead of a blank page.
Design → 03One Worker serves the game, the API and the database. It also serves the Discord Activity, so a single pushed tag ships both audiences.
Stack → 04The instrument reports its own limits. Thin rates carry 95% intervals, unmeasured days stay blank instead of reading as zero, and “too early to tell” is a verdict the dashboard can return.
Engineering → 05I logged a scheduling change, shipped it, then measured it. Finish rate moved from 81% to 91%. Read over a shorter window, the same change reports nothing.
Experiment →Cooking is my favorite hobby, and I wanted a game that taught people something about food they already eat. Like a lot of couples, my wife and I love to play daily guessing games together so combining that format with a food theme seemed like a promising idea. Initially I was inspired by the Wordle loop but major changes would need to be made to give the player a chance at guessing the daily dish. A word is an ordered list of letters, so green can mean right letter, right slot. A dish has no slots. Pad Thai and Bibimbap do not line up character by character in any way a player would care about.
I wrote down what the game had to do before I built any of it.
Readable at a glance
A player reads the board mid-round, between guesses. Anything needing a paragraph to interpret has already failed.
Effortlessly sharable
A result nobody can paste into a group chat gives the game no distribution. The grid is the entire marketing budget.
Narrowing enough to finish
A pool of uncountable dishes in someone's mind has to filter down to one after six guesses.
The game asks you to know things about food, and that knowledge is not handed out evenly. A player in Ohio and a player in Lagos open a round on Shakshuka from different places. I could have designed that away by scoring nothing except what the tiles hand you, and I kept it, because the gap is the part worth playing for.
Wordle can be played well without knowing the word. Letter frequency, position and elimination will carry you to answers you have never used in your life, which is a real skill and a skill at Wordle. Naming the Special is a skill at food. It draws on what you have eaten, what you have read, and what you can infer from a Tunisian breakfast sharing a pantry with a French summer stew.
Culinary knowledge also improves, which is what makes the unfairness survivable. Nobody takes every Special, and I would miss plenty myself if I were not the one booking the menu. A round you lose still ends with you knowing the dish, where it came from and what goes in it, so the next one starts further along. That is the deal the beat sheet makes: every miss buys information, and the information keeps working after the round is over.
An even start would have cost the game the thing that makes solving it feel like anything. I owed players a route from wherever they began instead, and that is what the beat sheet is for.
The first channel rewards deduction. The second concedes information as you fail. Together they let a player who knows Tunisian food and a player who does not both reach the same answer, at different speeds.
Channel one scores your guess on four attributes and on how its ingredients overlap the Special’s. Country carries the only middle state: the same country turns green, and a different country in the same region turns amber. That single warm signal is what makes triangulating across a map possible. Channel two is the clue ticket. Miss, and the kitchen slides the next one across the counter.
Today’s Special
0 of 6 guessesOrder something from the menu and the kitchen answers.
The menu (a fixed demo dish, not today’s answer)
The ingredient line reports which of your guess’s ingredients the Special also uses, and how many of the Special’s total that covers. It never lists the ones you have not found. Five of eight tells you three remain without telling you what they are.
A daily game eats a piece of authored content every 24 hours and never stops. Treating 364 dishes as 364 separate creative problems collapses around week eight, so I specified the dramatic order once instead: five beats, the same sequence for every dish, running from a map to a near-giveaway. Each beat states what its clue has to accomplish and what it must not give away, which is the difference between a brief and a blank page.
Broad geography
Narrows the map without naming a country.
Origin and history
The first beat with a voice, where the dish becomes a place rather than a puzzle token.
What makes it unmistakable
True of this dish and almost no other. The film, the street cart, the argument over who invented it, the thing that burns you. The beat people repeat to someone else afterwards.
A key ingredient or technique
Turns knowing about the dish into being able to name it.
Near-giveaway
Everything but the name. Missing here should feel unlucky, never unfair.
Clue 3 for Pierogi and clue 3 for Ceviche are the same job, and that sameness is what makes the pool extensible: the brief holds still while the subject changes. The beat sheet also fixes pacing, which is the part players feel. Beat 3 lands on the third miss, when someone needs a reason to stay in the round, and beat 5 arrives at the point where losing would otherwise feel arbitrary.
Margherita Pizza
A southern European classic, from a city in the shadow of a volcano. Gives the region and one image. Names nothing.
Born in Naples, where street ovens fed the working class for centuries. Country spent. The dish acquires a class and a century.
Named for a queen who visited in 1889 and picked the patriotic option. The beat that survives the round, and still not the answer.
Its three toppings, tomato, mozzarella and basil, mirror a national flag. Trades the story for the ingredients. Solvable from here.
The world’s most famous flatbread, best from a wood-fired oven. Everything but the name.
Beat 3 carries the heaviest load, so it is the beat the spec is strictest about. A player who loses the round should still come away with the queen and the flag, because that is the part they repeat to somebody at lunch. Beats 1 and 5 read as constraints more than as prose: one must not narrow the map too far, the other must leave nothing but the name. Judging whether a clue belongs on its beat takes about ten seconds, and that is the property I was after.
Each piece of interface takes the name it would have in a diner, and those names do work a tutorial would otherwise have to do. Nobody needs telling what a receipt is for.
| In the fiction | What it is |
|---|---|
| The Special | Today’s puzzle. One dish, every player, worldwide. |
| A clue ticket | The beat the kitchen hands you after a miss. |
| The check | End of round: result, streak, share. |
| Leftovers | The archive of days you missed. |
| Chef’s Choice | A random dish, no stakes, no streak. |
| A note from the kitchen | An announcement from me to players. |
🛎️ LUNCH SPECIAL
No. 29 · solved in 5
Lunch Special #29 — 5/6
⬜⬜🟩🟩 5/8🥄
⬜⬜🟩⬜ 1/8🥄
⬜🟩🟩⬜ 2/8🥄
🟨⬜🟩🟩 3/8🥄
🟩🟩🟩🟩 🛎️
lunchspecial.app
I specified that card before the feedback system and before the schema. A daily game with no shareable result has no distribution, so the shape of an emoji grid became a hard constraint on what the feedback was allowed to be: four tiles to a row, one pantry count, no dish names, and nothing that spoils the answer for whoever reads it over breakfast. For most people the card is the only part of this game they will see, which is a decent argument for designing it first.
Narrative design here is a production system. A fixed dramatic order keeps 1,820 clues consistent with each other, and the fiction absorbs most of the tutorial.
A spec that makes a dish a day routine is worth nothing if shipping that dish takes a week. So every choice below buys the same thing: the shortest distance one person can manage between having an idea and watching it get measured. One Cloudflare Worker serves the React bundle, answers the API and queries a SQLite database on the same platform. There is no second service to keep warm and no deploy that half succeeds. The whole game runs on four runtime dependencies, listed above.
| The goal it serves | Choice | What it bought |
|---|---|---|
| Idea to live in an evening | One Worker: React SPA, Hono API, D1 | One thing to deploy. Nothing to orchestrate between two services, and no state where half the release landed. |
| Playtesting that is not lying to me | @cloudflare/vite-plugin, worker in workerd locally |
My laptop and production run the same engine. No mock API to drift out of sync, so what I feel in a playtest is what players get. |
| A mistyped region must fail loudly | D1 with column-level CHECK constraints | A bad enum becomes a failed insert instead of a tile that never matches anything again and never says why. |
| No sign-up between a player and the game | No accounts. Player state in localStorage | Nothing personal to defend. The cost is real and I state it throughout: every audience figure here counts devices, not people. |
| A second audience without a second product | Discord Activity on the same deploy | The SDK sits behind a dynamic import and is fetched only inside the iframe, so web visitors never download it. All that ships to everyone is the check for the iframe. |
| Shipping without babysitting | CI/CD: GitHub Actions, CodeQL, Dependabot | A pushed tag runs the tests, the typecheck, the migrations and the deploy. Releasing on a weeknight is a five-minute decision. |
| Knowing whether the change worked | Analytics and an experiment log, in the admin panel | Measurement ships as part of the product. A daily game without it is one you tune by feel, which is how I got the difficulty wrong the first time. |
Request path
Only /api/* wakes the Worker. The board itself is a static file, so loading
the game costs nothing and only a guess spends compute.
Four runtime dependencies and one thing to deploy. I kept the stack unambitious so the time could go into the menu.
The game engine took a few days. Everything since has been the apparatus around it, because you operate a daily game rather than finish it. I built each station once the one before it started asking questions I could not answer. Authoring got a CMS once writing dishes as raw SQL stopped being funny, and that same panel took over the schedule when I lost track of what was booked. Measurement arrived the week I admitted I had no idea whether the puzzles were too easy.
One turn of the loop
Everything after the tag runs without me. Migrations go before the deploy, so the schema is never behind the code expecting it.
CI/CD · git tag v1.1.0 && git push origin v1.1.0
The Activity rides that same build. Discord does not host the code. It frames lunchspecial.app through its proxy under a URL mapping, so the tag that ships the website ships the Activity, and the only thing the browser fetches after it spots the iframe is Discord’s SDK. Two more workflows carry the routine load: tests and typecheck on every pull request, and a weekly security scan across the repository.
What I keep calling the admin panel is one page behind a password, and it has two halves. One is a CMS, a content management system: the place a dish and its five clues get written, validated and booked onto a date, doing the job for a menu that a newsroom tool does for articles. The other half is the dashboard that measures what players then do with them. Everything in the section above comes off that second half; the first is the one I would point at if you asked what I would bring to a content team.
The CMS authors a dish against a live preview of the real board, so I see the tiles a player will see before the dish is ever scheduled. It refuses to book a dish carrying fewer than three ingredients or anything other than five clues, which turns the beat sheet from a rule I have to remember into a rule the tool enforces. It books the calendar 30 days out, by hand or by an autofill that skips whatever has been served in the last 60 days.
None of that was necessary for one author. I built it as though somebody else would use it, because a schedule I keep in my head stops working the moment there are two of us, and because the constraints worth encoding are exactly the ones I would otherwise have to explain to every new writer.
The same panel posts notices. A note from the kitchen gets written, dated and aimed at an audience from the admin side, and putting one up is not a release. That is the difference between a game I operate and a game I redeploy.
One rule governs where they land. A notice appears on Today’s Special and nowhere else, never on a Leftover, a Chef’s Choice or a playtest, because those are side doors and somebody replaying a Thursday from three weeks ago did not come for an announcement. A first-timer gets the how-to first and the notice behind it. The card also drops in from above and bounces, where every other modal in the game rises from the bottom, so a player can tell a notice from their check before reading either. Afterwards the panel reports how many devices saw it, split by surface, which is the only way to find out whether a notice was read or merely posted.
A note from the kitchen1 of 2
Every one of them arrived through the suggest a dish form on your check, and each gets credited when it comes up as the Special. Keep them coming.
Where a notice appears
The bottom three are side doors, and a notice waiting behind one of them would interrupt somebody who came for a different reason.
The loop runs the other way too. After finishing a round, a button on the bottom of the modal shows "suggest a dish for the menu" that prompts for a dish to be added. Players can give the name, a country if you know it, and a note.
Suggestions land in an inbox on the admin side where each one is either cleared or promoted, and promoting it opens the dish editor with the name and country already in the fields, at which point it has to clear the same bar as everything else on the menu.
· · · the rest of the check · · ·
Next Special in 6:12:44
Off a customer’s ticket
A regular asked for Zeppole. Yours could be next.
What the credit changes
One flag, read in exactly one place. A dish that came in through the form is judged on the same terms as the rest of the menu.
Where the seal sits is a height decision rather than a layout preference. The check is the tallest card in the game on a 375px screen, so the credit goes at the very foot of it, directly on top of the suggest button, instead of up beside the dish name where it would read as a label on the answer. It is square rather than tilted, because a tilt reads as a sticker and then demands padding on four corners to keep them off the text around it. And it carries the section break itself: the dashed rule the suggest form draws above itself is dropped whenever the seal is there, so crediting a dish costs the card one short band rather than a separator and a seal. Mustard, not cherry, because nothing on that card is allowed to outrank the verdict.
Whether the credit brings a player back on the day their dish runs, I do not know. It is not instrumented, and the honest version of this page says so rather than assuming the flattering answer.
Queries need a database. Arithmetic does not. Every number on this page comes out of a pure function that takes rows and returns a shape, which is why 333 tests sit behind them and why I trust a chart I have never eyeballed. The public endpoint this page reads runs the same folds as my private dashboard, so the two cannot drift into quoting different numbers at each other.
The client keeps the round’s bookkeeping and asks for the reveal after game over, the way Wordle does. Anyone determined can read the answer out of a network tab. Closing that needs server-side sessions, and therefore accounts. The cheat costs one player their own round; the account system would cost every player a sign-up.
A daily game is an operation rather than a launch. I built the stations in the order the game demanded them.
Everything below comes out of production the moment you load this page. The audience is small and self-selected: 939 rounds from 312 devices, of whom 58 came back on a second day. Read it as a method and a first result rather than a finding about daily games in general.
At this size the honest problem is that a rate off 12 completions and a rate off 400 look identical printed. So any rate I quote from fewer than 30 observations carries its 95% interval, and nothing else does, because an interval on every number teaches you to skip them. The intervals are Wilson rather than the textbook formula, because this game lives exactly where the textbook one fails: tiny samples, and rates pinned near 0 or 1, where it returns an interval of zero width. The most confident thing a small dashboard can print is usually the most wrong.
Where the counter reaches
one country per device, all time
Most players first
10+ devices 2–9 1
Guesses needed to solve
684 solved rounds · 140 more ran out of guesses
The full six-guess budget gets used and a real share of rounds end unsolved, which is the shape I wanted and did not have at launch. See the third finding below.
Which entrance players use
rounds started, by mode
Where players drop off
devices, pooled over the days arrivals have been measured
Every stage counts devices rather than rounds, which is the only reason this is allowed to be a funnel. One player doing the Special plus three Leftovers is one arrival and four starts, and stacked as stages those numbers would grow as they descend.
A round counts as started on the player’s first submitted guess. Opening a page is not playing a game, and I still think that is the correct definition. It also meant that everyone who arrived and never guessed was invisible to me. Games started was the widest number I had, with nothing above it, so a change that doubled interest while halving conversion would have looked identical to no change at all.
One row per device per ET day closed it, and the Arrived row above is what that bought. The rows underneath it changed meaning at the same time: the first row counts people and every row below counts games, so dividing one straight into the other reports play rates over 100%. Both rows name their unit for that reason. I also had to give the admin panel a way to delete its own play-testing, because at tens of rounds a day one person reloading the game to check a change writes arrival rows and nothing else, and those rows read as pure bounce in the funnel I had built the week before.
Games played, running total
every mode together, since the first recorded round
A cumulative curve can only rise, so the fact it goes up carries no information. The bend does. The dashed line is the same run at one constant pace, fitted by least squares, and the curve pulling above it means the game is gaining. Without that reference every cumulative chart looks like success.
The archive became a third of all play
I built Leftovers as a courtesy for anyone who found the game late. Together with Chef’s Choice it now accounts for 35% of every round played, and 93% of those rounds run to the end, against 85% for the Special itself.
So: I stopped treating the archive as an accessory. It now unlocks the moment today’s check is settled rather than sitting behind a menu, and the dish report in my dashboard counts a Leftover as a genuine first attempt at that dish, which tripled the sample behind every per-dish difficulty read.
The Discord Activity out-recruits the website
Discord reaches 174 devices against the web’s 139, off 363 rounds against 576. More people, each playing less. The embed recruits and the website is where the habit forms.
So: I built the loop that suits recruiting rather than the one that suits retention. A player’s first guess posts one message into the channel they launched from, showing their board and a button to play; every later guess edits that same message, so a whole round costs the channel one post. I am watching whether it converts, and it is the next thing I will log as an experiment.
The menu was too easy, and the data said so before I did
Mean solve sat at 2.8 guesses. The sixth guess went unused and the last two clue beats went unread, which meant I had written content nobody reached. The pool skewed toward globally famous dishes, good for a launch and wrong by week four. I added 60 dishes chosen to close the gaps. Mean is now 3.2, the budget gets spent, and 17% of finished rounds end unsolved.
The instrument told me the game was too easy and that my courtesy feature was carrying a third of the play. Neither was visible from inside the design.
Every chart above says what happened. None of them says whether I caused it. So the database also holds a change log: one row per deliberate change, carrying what I expected and which metric it was meant to move, both written down before the change went out. The dashboard splits the daily series at the ship date and compares the halves.
🛎️ THE CHECK
Experiment log · shipped 26 Jul 2026
MOVED UP
intervals do not overlap
I had picked the schedule for variety by instinct until this day. From here I booked it against the menu-mix panel, which compares the region, course, protein and temperature ratios of what I have served against the pool those dishes are drawn from. More of the board became guessable from the tiles alone, and more players reached the end.
The fixed 14 days applies to the after period, which is the half I could have gamed by stopping on a good week. The before period is 8 days because the game was 8 days old when I shipped the change, so that is the entire history that existed. An asymmetry I picked would be a problem. This one is a ceiling, and the smaller sample sits on the side that makes the result harder to claim rather than easier: the before period’s interval is the wider of the two, and the verdict still needed them not to overlap.
Read at 7 days instead of 14, the same comparison returns no change: 84% against 89%, with intervals that overlap. Nothing about the game differs between those two readings, only how much of the surrounding variation each window swallows. Choosing the window after seeing the data means choosing the answer.
Finish rate was the metric because the arrivals beacon in the section above had made it readable. Naming a metric up front only works when you can already measure the thing you are about to name.
I claim a change once it separates from ordinary variation, and the system tells me when it cannot.
Lunch Special is one designer, one database and a few hundred players. I am not going to pretend that is a shipped title. Three of the habits behind it are the reason I built it this way, and those are what I would bring to a team.
A beat sheet is a content pipeline
The five beats exist so that a dish is a repeatable job rather than an inspiration. On a bigger game the same spec is what keeps six writers on one voice, and what lets you cost a content month instead of guessing at it. I built the CMS to enforce it rather than trusting anyone to remember it, and that includes me.
A named metric settles an argument that seniority would otherwise settle
Writing down the metric and the window before shipping costs ten minutes and removes the move where everyone reads whichever number happened to move. On a team that is worth more than it is solo, because solo the only person I can fool is me.
An instrument that cannot say “I don’t know” will say something else
Unmeasured days stay blank, thin rates carry their intervals, and “too early to tell” is a verdict with a date attached. The failure this prevents scales up rather than down: a telemetry gap read as a player collapse is an expensive mistake with a publisher in the room.
It does not show me working with other people, because there were none. I have never had to defend a pre-registered metric to a producer who needs the number to move by Friday, or hand this spec to a writer whose instincts differ from mine and find out which parts of it were load-bearing and which were just my taste. Those are the parts I want next, and they are the reason I am looking at studios rather than at another solo project.
The scale is small and the habits are not: write the rules down, build the tool that enforces them, and measure the change you claimed you were making.