
When Your AI Change 'Works' But Success Drops 20 Points
Prompt and model tweaks that pass every smoke test can quietly wreck real outcomes. Here's how to evaluate agents on results instead of vibes.
Blog
Practical writing on productivity, AI, and building software.

Prompt and model tweaks that pass every smoke test can quietly wreck real outcomes. Here's how to evaluate agents on results instead of vibes.

Public benchmarks won't tell you if a model can triage your inbox or reconcile your spreadsheet. Here's how to build a private, brutally specific eval set in an afternoon.

A chatbot returns a response. An agent decides and executes a sequence. Confusing the two is why so many 'agentic' rollouts quietly stall out.

Picking a single LLM for every task leaves capability and money on the table. Route by task instead: cheap triage, strong drafting, dedicated verification.

Email isn't a list of messages — it's an unindexed database you never designed. Here's how to query it in natural language without hallucinating your way into a bad reply.

Medical research on why some experts resist bad AI advice — and how to turn those findings into a practical framework for judging LLM reliability in your own work.

Email, calendar, and file storage are already chunked, timestamped, and permission-aware. Treating them as your retrieval corpus beats dumping documents into a vector database.

A practical walkthrough for building a 40-example eval set in a spreadsheet — plus the LLM-as-judge traps that make most accuracy numbers meaningless.

Agents that log in as you aren't a feature — they're a security architecture failure. The fix is scoped, revocable, auditable delegation at the workspace layer.

A concrete walkthrough of building a Tamaton agent that triages inbox, drafts replies with real context, and books follow-ups — with a human holding the approval button.

Only ~13% of IT orgs have sanctioned AI agents — but unsanctioned ones are already running on employee credentials with no scopes, no logs, and no way to revoke access.

Long, harmless context isn't neutral. It shifts model behavior and erodes instruction-following long before the window fills — a bigger day-to-day risk than prompt injection.
Get started
Claim your address before someone else does — free to start, with an AI-native inbox built in.