Projects
-
Model Arena
Ten small tasks; three or four models race head-to-head. Scoring is deterministic — correct → cheapest → fastest — so the winner is arithmetic, not vibes. Replays are recorded and free; live runs stream real tokens behind a bot-check.
typical race$0.0009 (July 2026)full live sweep~$0.02One stream feeds every lane at once; scoring is a fixed rule; live spend stops at $0.50 a day.
Next: visitor-submitted challenges, moderated, with the same scoring discipline.
-
Guess Which AI
A daily parlor game: one odd prompt, three anonymous answers — on some days a fourth written by me. Spot who wrote what. The reveal prints each answer’s model, capture date, and exact price.
cost per answer$0.0000099–0.000606 (July 2026)my answers$0.00 + ~4 minAnswers recorded once, first take kept; sets chosen by rules written down in advance; nothing runs server-side.
Next: a season-two roster swap once the current models’ tells get famous.
-
Token Receipt
The tokenizer under this site’s hero. Exact OpenAI counts computed in your browser via gpt-tokenizer (o200k); Claude counts shown as labeled estimates because Claude’s tokenizer is private. Nothing you type leaves the page.
runs onyour devicecost to run$0.000000Client-side compute, a lazy-loaded tokenizer bundle, estimates labeled as estimates.
Next: side-by-side cost diffing for two prompts.
-
Linework
Daily puzzle games for writing craft — the shape and ritual of the NYT word games, but the skill being exercised is prose. Started on a bet that as AI drafts more of everyone’s words, deliberate practice for human writing gets more valuable, not less.
excerpts prepared for testing100+ (July 2026)statusin progress — three apps nearly done (July 2026)Next: finish the first three apps and grow the excerpt bank to launch depth.
-
Pricing Optimizer
A simulate-and-compare tool for pricing decisions. The chart wiring came out genuinely sleek. The demand data underneath was generated, so the answers couldn’t mean much — and that’s why it stopped.
statusbuilt late 2025 · discontinued early 2026Next: a rebuilt “Pricing Explorer” — same domain, aimed at exploring and comparing pricing models rather than optimizing one.
-
SEC Narrative Drift
Diffs the story a company tells between SEC filings — which risk paragraphs appeared, vanished, or quietly changed adjectives.
Risk 1A sections cleaned50+ S&P 500 companies, 2016–2025statusDec 2025 – Mar 2026 · paused, unfinishedNext: parked. The extraction practice was the real yield; the year-over-year diffs were mostly noise, and that lesson cost three months.
-
Response Option Fit
Checks whether a survey’s answer options actually fit its question — how survey data goes wrong before anyone answers.
statusa design study · spring 2026Next: none planned — it did its job as a study.