🛎️ Lunch Special Play it

Project breakdown · 2026

Lunch Special

Guess the day’s dish.

Jacob Poteet Senior Technical Narrative Designer. Solo creator of Lunch Special.

Every guess is a real dish, and the kitchen tells you what that dish shares with the Special. I designed and implemented the gameplay loop, the CMS for the menu of 364 dishes, the telemetry underneath, and the experiment log.

Serving Special No. — · next one in

939 Games played across three modes, two surfaces
364 Dishes written 1,820 clues, five to a dish
152 Results shared emoji grids posted out of the game
19% Came back another day 58 of 312 devices, no accounts

Read live from the production database when this page loads.

The short version

01

What I set out to make

goals

Cooking is my favorite hobby, and I wanted a game that taught people something about food they already eat. Like a lot of couples, my wife and I love to play daily guessing games together so combining that format with a food theme seemed like a promising idea. Initially I was inspired by the Wordle loop but major changes would need to be made to give the player a chance at guessing the daily dish. A word is an ordered list of letters, so green can mean right letter, right slot. A dish has no slots. Pad Thai and Bibimbap do not line up character by character in any way a player would care about.

I wrote down what the game had to do before I built any of it.

Goal 01

Readable at a glance

A player reads the board mid-round, between guesses. Anything needing a paragraph to interpret has already failed.

Goal 02

Effortlessly sharable

A result nobody can paste into a group chat gives the game no distribution. The grid is the entire marketing budget.

Goal 03

Narrowing enough to finish

A pool of uncountable dishes in someone's mind has to filter down to one after six guesses.

I left the unfairness in

The game asks you to know things about food, and that knowledge is not handed out evenly. A player in Ohio and a player in Lagos open a round on Shakshuka from different places. I could have designed that away by scoring nothing except what the tiles hand you, and I kept it, because the gap is the part worth playing for.

Wordle can be played well without knowing the word. Letter frequency, position and elimination will carry you to answers you have never used in your life, which is a real skill and a skill at Wordle. Naming the Special is a skill at food. It draws on what you have eaten, what you have read, and what you can infer from a Tunisian breakfast sharing a pantry with a French summer stew.

Culinary knowledge also improves, which is what makes the unfairness survivable. Nobody takes every Special, and I would miss plenty myself if I were not the one booking the menu. A round you lose still ends with you knowing the dish, where it came from and what goes in it, so the next one starts further along. That is the deal the beat sheet makes: every miss buys information, and the information keeps working after the round is over.

Takeaway

An even start would have cost the game the thing that makes solving it feel like anything. I owed players a route from wherever they began instead, and that is what the beat sheet is for.

02

How I designed it

feedback · clues · fiction

Two channels, running at once

The first channel rewards deduction. The second concedes information as you fail. Together they let a player who knows Tunisian food and a player who does not both reach the same answer, at different speeds.

Channel one scores your guess on four attributes and on how its ingredients overlap the Special’s. Country carries the only middle state: the same country turns green, and a different country in the same region turns amber. That single warm signal is what makes triangulating across a map possible. Channel two is the clue ticket. Miss, and the kitchen slides the next one across the counter.

Today’s Special

0 of 6 guesses

Order something from the menu and the kitchen answers.

The menu (a fixed demo dish, not today’s answer)

The ingredient line reports which of your guess’s ingredients the Special also uses, and how many of the Special’s total that covers. It never lists the ones you have not found. Five of eight tells you three remain without telling you what they are.

One beat sheet specifies every clue

A daily game eats a piece of authored content every 24 hours and never stops. Treating 364 dishes as 364 separate creative problems collapses around week eight, so I specified the dramatic order once instead: five beats, the same sequence for every dish, running from a map to a near-giveaway. Each beat states what its clue has to accomplish and what it must not give away, which is the difference between a brief and a blank page.

1

Broad geography

Narrows the map without naming a country.

2

Origin and history

The first beat with a voice, where the dish becomes a place rather than a puzzle token.

3

What makes it unmistakable

True of this dish and almost no other. The film, the street cart, the argument over who invented it, the thing that burns you. The beat people repeat to someone else afterwards.

4

A key ingredient or technique

Turns knowing about the dish into being able to name it.

5

Near-giveaway

Everything but the name. Missing here should feel unlucky, never unfair.

Clue 3 for Pierogi and clue 3 for Ceviche are the same job, and that sameness is what makes the pool extensible: the brief holds still while the subject changes. The beat sheet also fixes pacing, which is the part players feel. Beat 3 lands on the third miss, when someone needs a reason to stay in the round, and beat 5 arrives at the point where losing would otherwise feel arbitrary.

Margherita Pizza one dish against the spec

1

A southern European classic, from a city in the shadow of a volcano. Gives the region and one image. Names nothing.

2

Born in Naples, where street ovens fed the working class for centuries. Country spent. The dish acquires a class and a century.

3

Named for a queen who visited in 1889 and picked the patriotic option. The beat that survives the round, and still not the answer.

4

Its three toppings, tomato, mozzarella and basil, mirror a national flag. Trades the story for the ingredients. Solvable from here.

5

The world’s most famous flatbread, best from a wood-fired oven. Everything but the name.

Beat 3 carries the heaviest load, so it is the beat the spec is strictest about. A player who loses the round should still come away with the queen and the flag, because that is the part they repeat to somebody at lunch. Beats 1 and 5 read as constraints more than as prose: one must not narrow the map too far, the other must leave nothing but the name. Judging whether a clue belongs on its beat takes about ten seconds, and that is the property I was after.

The fiction carries the interface

Each piece of interface takes the name it would have in a diner, and those names do work a tutorial would otherwise have to do. Nobody needs telling what a receipt is for.

In the fiction What it is
The SpecialToday’s puzzle. One dish, every player, worldwide.
A clue ticketThe beat the kitchen hands you after a miss.
The checkEnd of round: result, streak, share.
LeftoversThe archive of days you missed.
Chef’s ChoiceA random dish, no stakes, no streak.
A note from the kitchenAn announcement from me to players.

🛎️ LUNCH SPECIAL

No. 29 · solved in 5

Lunch Special #29 — 5/6
⬜⬜🟩🟩 5/8🥄
⬜⬜🟩⬜ 1/8🥄
⬜🟩🟩⬜ 2/8🥄
🟨⬜🟩🟩 3/8🥄
🟩🟩🟩🟩 🛎️

lunchspecial.app

I specified that card before the feedback system and before the schema. A daily game with no shareable result has no distribution, so the shape of an emoji grid became a hard constraint on what the feedback was allowed to be: four tiles to a row, one pantry count, no dish names, and nothing that spoils the answer for whoever reads it over breakfast. For most people the card is the only part of this game they will see, which is a decent argument for designing it first.

Takeaway

Narrative design here is a production system. A fixed dramatic order keeps 1,820 clues consistent with each other, and the fiction absorbs most of the tutorial.

03

The stack, and what each choice bought

hono · react · react-dom · discord sdk

A spec that makes a dish a day routine is worth nothing if shipping that dish takes a week. So every choice below buys the same thing: the shortest distance one person can manage between having an idea and watching it get measured. One Cloudflare Worker serves the React bundle, answers the API and queries a SQLite database on the same platform. There is no second service to keep warm and no deploy that half succeeds. The whole game runs on four runtime dependencies, listed above.

The goal it serves Choice What it bought
Idea to live in an evening One Worker: React SPA, Hono API, D1 One thing to deploy. Nothing to orchestrate between two services, and no state where half the release landed.
Playtesting that is not lying to me @cloudflare/vite-plugin, worker in workerd locally My laptop and production run the same engine. No mock API to drift out of sync, so what I feel in a playtest is what players get.
A mistyped region must fail loudly D1 with column-level CHECK constraints A bad enum becomes a failed insert instead of a tile that never matches anything again and never says why.
No sign-up between a player and the game No accounts. Player state in localStorage Nothing personal to defend. The cost is real and I state it throughout: every audience figure here counts devices, not people.
A second audience without a second product Discord Activity on the same deploy The SDK sits behind a dynamic import and is fetched only inside the iframe, so web visitors never download it. All that ships to everyone is the check for the iframe.
Shipping without babysitting CI/CD: GitHub Actions, CodeQL, Dependabot A pushed tag runs the tests, the typecheck, the migrations and the deploy. Releasing on a weeknight is a five-minute decision.
Knowing whether the change worked Analytics and an experiment log, in the admin panel Measurement ships as part of the product. A daily game without it is one you tune by feel, which is how I got the difficulty wrong the first time.

Request path

Browser lunchspecial.app or a Discord iframe ONE WORKER, AT THE EDGE Static assets SPA, fonts, art no Worker wake Hono API /api/* only guesses, beacons HMAC auth stateless cookie admin + preview D1 (SQLite) dishes, clues, schedule, telemetry, experiments

Only /api/* wakes the Worker. The board itself is a static file, so loading the game costs nothing and only a guess spends compute.

Takeaway

Four runtime dependencies and one thing to deploy. I kept the stack unambitious so the time could go into the menu.

04

How it was engineered

333 tests · 23 files

The game engine took a few days. Everything since has been the apparatus around it, because you operate a daily game rather than finish it. I built each station once the one before it started asking questions I could not answer. Authoring got a CMS once writing dishes as raw SQL stopped being funny, and that same panel took over the schedule when I lost track of what was booked. Measurement arrived the week I admitted I had no idea whether the puzzles were too easy.

One turn of the loop

ONE TURN OF THE LOOP 01 Author dish + 5 clues 02 Schedule 30 days out 03 Play web + Discord 04 Measure 4 beacons 05 Hypothesise metric first 06 Ship tag → live

A release starts with a tag

Everything after the tag runs without me. Migrations go before the deploy, so the schema is never behind the code expecting it.

CI/CD · git tag v1.1.0 && git push origin v1.1.0

Trigger Push a v* tag deploy.yml
Gate 1 Tests 333 across 23 files
Gate 2 Typecheck tsc -b, 3 projects
Schema Migrate D1 additive, idempotent
Release Deploy the Worker one build
🌐 lunchspecial.app the public game
🎮 The Discord Activity same build, framed through Discord’s proxy

The Activity rides that same build. Discord does not host the code. It frames lunchspecial.app through its proxy under a URL mapping, so the tag that ships the website ships the Activity, and the only thing the browser fetches after it spots the iframe is Discord’s SDK. Two more workflows carry the routine load: tests and typecheck on every pull request, and a weekly security scan across the repository.

The tool I built for the writer, who happens to be me

What I keep calling the admin panel is one page behind a password, and it has two halves. One is a CMS, a content management system: the place a dish and its five clues get written, validated and booked onto a date, doing the job for a menu that a newsroom tool does for articles. The other half is the dashboard that measures what players then do with them. Everything in the section above comes off that second half; the first is the one I would point at if you asked what I would bring to a content team.

The CMS authors a dish against a live preview of the real board, so I see the tiles a player will see before the dish is ever scheduled. It refuses to book a dish carrying fewer than three ingredients or anything other than five clues, which turns the beat sheet from a rule I have to remember into a rule the tool enforces. It books the calendar 30 days out, by hand or by an autofill that skips whatever has been served in the last 60 days.

None of that was necessary for one author. I built it as though somebody else would use it, because a schedule I keep in my head stops working the moment there are two of us, and because the constraints worth encoding are exactly the ones I would otherwise have to explain to every new writer.

Talking to players without shipping a build

The same panel posts notices. A note from the kitchen gets written, dated and aimed at an audience from the admin side, and putting one up is not a release. That is the difference between a game I operate and a game I redeploy.

One rule governs where they land. A notice appears on Today’s Special and nowhere else, never on a Leftover, a Chef’s Choice or a playtest, because those are side doors and somebody replaying a Thursday from three weeks ago did not come for an announcement. A first-timer gets the how-to first and the notice behind it. The card also drops in from above and bounces, where every other modal in the game rises from the bottom, so a player can tell a notice from their check before reading either. Afterwards the panel reports how many devices saw it, split by surface, which is the only way to find out whether a notice was read or merely posted.

A note from the kitchen1 of 2

Thirteen new dishes on the menu

Every one of them arrived through the suggest a dish form on your check, and each gets credited when it comes up as the Special. Keep them coming.

Where a notice appears

  • Today’s Special
  • Leftovers
  • Chef’s Choice
  • Preview & playtest

The bottom three are side doors, and a notice waiting behind one of them would interrupt somebody who came for a different reason.

Players add to the menu

The loop runs the other way too. After finishing a round, a button on the bottom of the modal shows "suggest a dish for the menu" that prompts for a dish to be added. Players can give the name, a country if you know it, and a note.

Suggestions land in an inbox on the admin side where each one is either cleared or promoted, and promoting it opens the dish editor with the name and country already in the fields, at which point it has to clear the same bar as everything else on the menu.

· · · the rest of the check · · ·

Next Special in 6:12:44

Off a customer’s ticket

A regular asked for Zeppole. Yours could be next.

🍽️ Suggest a dish for the menu

What the credit changes

  • The seal on the check
  • Which day it is served
  • The odds in Chef’s Choice
  • The tiles and the clues
  • Every figure on the dashboard

One flag, read in exactly one place. A dish that came in through the form is judged on the same terms as the rest of the menu.

Where the seal sits is a height decision rather than a layout preference. The check is the tallest card in the game on a 375px screen, so the credit goes at the very foot of it, directly on top of the suggest button, instead of up beside the dish name where it would read as a label on the answer. It is square rather than tilted, because a tilt reads as a sticker and then demands padding on four corners to keep them off the text around it. And it carries the section break itself: the dashed rule the suggest form draws above itself is dropped whenever the seal is there, so crediting a dish costs the card one short band rather than a separator and a seal. Mustard, not cherry, because nothing on that card is allowed to outrank the verdict.

Whether the credit brings a player back on the day their dish runs, I do not know. It is not instrumented, and the honest version of this page says so rather than assuming the flattering answer.

The arithmetic is separated from the database

Queries need a database. Arithmetic does not. Every number on this page comes out of a pure function that takes rows and returns a shape, which is why 333 tests sit behind them and why I trust a chart I have never eyeballed. The public endpoint this page reads runs the same folds as my private dashboard, so the two cannot drift into quoting different numbers at each other.

One trade-off, made with my eyes open

The client keeps the round’s bookkeeping and asks for the reveal after game over, the way Wordle does. Anyone determined can read the answer out of a network tab. Closing that needs server-side sessions, and therefore accounts. The cheat costs one player their own round; the account system would cost every player a sign-up.

Takeaway

A daily game is an operation rather than a launch. I built the stations in the order the game demanded them.

05

What the numbers say

read live

Everything below comes out of production the moment you load this page. The audience is small and self-selected: 939 rounds from 312 devices, of whom 58 came back on a second day. Read it as a method and a first result rather than a finding about daily games in general.

At this size the honest problem is that a rate off 12 completions and a rate off 400 look identical printed. So any rate I quote from fewer than 30 observations carries its 95% interval, and nothing else does, because an interval on every number teaches you to skip them. The intervals are Wilson rather than the textbook formula, because this game lives exactly where the textbook one fails: tiny samples, and rates pinned near 0 or 1, where it returns an interval of zero width. The most confident thing a small dashboard can print is usually the most wrong.

Where the counter reaches

one country per device, all time

Most players first

10+ devices 2–9 1

Guesses needed to solve

684 solved rounds · 140 more ran out of guesses

Solved Ran out of guesses

The full six-guess budget gets used and a real share of rounds end unsolved, which is the shape I wanted and did not have at launch. See the third finding below.

Which entrance players use

rounds started, by mode

Where players drop off

devices, pooled over the days arrivals have been measured

Every stage counts devices rather than rounds, which is the only reason this is allowed to be a funnel. One player doing the Special plus three Leftovers is one arrival and four starts, and stacked as stages those numbers would grow as they descend.

That top row cost me a release, and I had the definition right

A round counts as started on the player’s first submitted guess. Opening a page is not playing a game, and I still think that is the correct definition. It also meant that everyone who arrived and never guessed was invisible to me. Games started was the widest number I had, with nothing above it, so a change that doubled interest while halving conversion would have looked identical to no change at all.

One row per device per ET day closed it, and the Arrived row above is what that bought. The rows underneath it changed meaning at the same time: the first row counts people and every row below counts games, so dividing one straight into the other reports play rates over 100%. Both rows name their unit for that reason. I also had to give the admin panel a way to delete its own play-testing, because at tens of rounds a day one person reloading the game to check a change writes arrival rows and nothing else, and those rows read as pure bounce in the funnel I had built the week before.

Games played, running total

every mode together, since the first recorded round

Running total Constant-pace reference

A cumulative curve can only rise, so the fact it goes up carries no information. The bend does. The dashed line is the same run at one constant pace, fitted by least squares, and the curve pulling above it means the game is gaining. Without that reference every cumulative chart looks like success.

Three things I did not expect

Finding 01

The archive became a third of all play

I built Leftovers as a courtesy for anyone who found the game late. Together with Chef’s Choice it now accounts for 35% of every round played, and 93% of those rounds run to the end, against 85% for the Special itself.

So: I stopped treating the archive as an accessory. It now unlocks the moment today’s check is settled rather than sitting behind a menu, and the dish report in my dashboard counts a Leftover as a genuine first attempt at that dish, which tripled the sample behind every per-dish difficulty read.

Finding 02

The Discord Activity out-recruits the website

Discord reaches 174 devices against the web’s 139, off 363 rounds against 576. More people, each playing less. The embed recruits and the website is where the habit forms.

So: I built the loop that suits recruiting rather than the one that suits retention. A player’s first guess posts one message into the channel they launched from, showing their board and a button to play; every later guess edits that same message, so a whole round costs the channel one post. I am watching whether it converts, and it is the next thing I will log as an experiment.

Finding 03

The menu was too easy, and the data said so before I did

Mean solve sat at 2.8 guesses. The sixth guess went unused and the last two clue beats went unread, which meant I had written content nobody reached. The pool skewed toward globally famous dishes, good for a launch and wrong by week four. I added 60 dishes chosen to close the gaps. Mean is now 3.2, the budget gets spent, and 17% of finished rounds end unsolved.

Takeaway

The instrument told me the game was too easy and that my courtesy feature was carrying a third of the play. Neither was visible from inside the design.

06

One change, logged before I shipped it

verdict · 14-day window

Every chart above says what happened. None of them says whether I caused it. So the database also holds a change log: one row per deliberate change, carrying what I expected and which metric it was meant to move, both written down before the change went out. The dashboard splits the daily series at the ship date and compares the halves.

🛎️ THE CHECK

Experiment log · shipped 26 Jul 2026

ChangeSchedule the menu from the mix panel, not by feel
Metric, named up frontFinish rate
Window, fixed in advance14 days after the ship date
Before period8 days: everything the game had
Before · 8 days 81%
95% CI 76–85% · 312 rounds
After · 14 days 91%
95% CI 88–93% · 404 rounds

MOVED UP

intervals do not overlap

I had picked the schedule for variety by instinct until this day. From here I booked it against the menu-mix panel, which compares the region, course, protein and temperature ratios of what I have served against the pool those dishes are drawn from. More of the board became guessable from the tiles alone, and more players reached the end.

The two windows are not the same length, and that is not a choice

The fixed 14 days applies to the after period, which is the half I could have gamed by stopping on a good week. The before period is 8 days because the game was 8 days old when I shipped the change, so that is the entire history that existed. An asymmetry I picked would be a problem. This one is a ceiling, and the smaller sample sits on the side that makes the result harder to claim rather than easier: the before period’s interval is the wider of the two, and the verdict still needed them not to overlap.

Why the after window gets fixed before the change ships

Read at 7 days instead of 14, the same comparison returns no change: 84% against 89%, with intervals that overlap. Nothing about the game differs between those two readings, only how much of the surrounding variation each window swallows. Choosing the window after seeing the data means choosing the answer.

The metric is named first Picked afterwards, it is whichever of the eight moved most. Reading a different one than the change was logged against prints a warning.
Rates are pooled, never averaged Total solved over total completed. Averaging daily percentages lets a one-round Tuesday outvote a whole weekend.
Counts and rates get different tests A count’s noise is day-to-day spread. A rate’s is binomial. Judging a share rate by how much traffic wobbled answers the wrong question.
“Too early to tell” is a verdict At this volume it is the honest answer for a fortnight, and it arrives with an estimate of how many more days the change needs.

Finish rate was the metric because the arrivals beacon in the section above had made it readable. Naming a metric up front only works when you can already measure the thing you are about to name.

Takeaway

I claim a change once it separates from ordinary variation, and the system tells me when it cannot.

07

What of this transfers to a bigger game

the honest version

Lunch Special is one designer, one database and a few hundred players. I am not going to pretend that is a shipped title. Three of the habits behind it are the reason I built it this way, and those are what I would bring to a team.

Transfers 01

A beat sheet is a content pipeline

The five beats exist so that a dish is a repeatable job rather than an inspiration. On a bigger game the same spec is what keeps six writers on one voice, and what lets you cost a content month instead of guessing at it. I built the CMS to enforce it rather than trusting anyone to remember it, and that includes me.

Transfers 02

A named metric settles an argument that seniority would otherwise settle

Writing down the metric and the window before shipping costs ten minutes and removes the move where everyone reads whichever number happened to move. On a team that is worth more than it is solo, because solo the only person I can fool is me.

Transfers 03

An instrument that cannot say “I don’t know” will say something else

Unmeasured days stay blank, thin rates carry their intervals, and “too early to tell” is a verdict with a date attached. The failure this prevents scales up rather than down: a telemetry gap read as a player collapse is an expensive mistake with a publisher in the room.

What this project does not show you

It does not show me working with other people, because there were none. I have never had to defend a pre-registered metric to a producer who needs the number to move by Friday, or hand this spec to a writer whose instincts differ from mine and find out which parts of it were load-bearing and which were just my taste. Those are the parts I want next, and they are the reason I am looking at studios rather than at another solo project.

Takeaway

The scale is small and the habits are not: write the rules down, build the tool that enforces them, and measure the change you claimed you were making.