max_tokens
Notes on LLM agents, inference, data systems, AI visibility, vibe coding, and building things on the internet.
- 2026-07-17 Setting up OMP in one evening: model roles, global skills, an advisor, and memory
- 2026-07-16 The agent passed the eval. I still wouldn't ship it
- 2026-07-09 What one headless agent run actually costs
- 2026-07-01 Orca vs Herdr: task isolation or live terminal control
- 2026-06-30 Loop engineering is not prompt engineering with a timer
- 2026-06-28 Turn count as a product metric
- 2026-06-19 Where prompt caching quietly misses: TTL, prefix order, and what invalidates the cache
- 2026-06-18 Not every request deserves an agent
- 2026-06-15 Graphify and MemPalace: an agent needs a project map and a decision history, not 'memory'
- 2026-06-14 Herdr: a control room for agents in the terminal
- 2026-06-13 Spec-Driven Development: from chaotic AI coding to an engineering process
- 2026-06-12 Why we chose Temporal to manage AI agents instead of Celery
- 2026-06-11 The prompt is 5% of the work: context engineering for a production LLM agent