
Your Inbox Is an API Now: Designing Work for Agent Readers
Agents don't click, they parse. Here's how to name, thread, permission and store your work so AI agents in the workplace can actually use it.

Agents don't click, they parse. Here's how to name, thread, permission and store your work so AI agents in the workplace can actually use it.

Public benchmarks won't tell you which model can triage your email. Here's how to build a small, honest eval set from your own inbox, docs and sheets — and what ours revealed.

Agents aren't magic buttons. Give them a scoped role, a 30-day ramp, real feedback loops, and review cycles — the same things you'd give a competent new teammate.

The bottleneck in production agent systems isn't autonomy — it's the escalation contract. Here's how to design triggers, handoff packages, and ownership rules that actually hold.

A practical test kit for knowledge workers: perturb the premises, audit the citation chain, and check whether the answer moves for the right reasons.

An analysis of 880,000+ texts found model-assisted writing converges on one style while meaning survives. Here's how to keep your voice with grounding, retrieval and better evals.

The real attack surface for AI agents isn't a clever chat jailbreak — it's the forwarded email, shared sheet, or PDF your agent reads with your permissions.

A concrete look at building RAG over your own email, docs, and drive — and why retrieval quality, not model size, decides whether AI knowledge work is actually useful.

Coding demos are graded by a test suite. Email is graded by your boss, your customer, and your calendar. Here's why the inbox is the hardest honest test of an AI agent.

Reasoning and planning demos are easy. Proving a multi-step agent actually finished the job — correctly, safely, once — is the unsolved part. Here's how to measure it.

"Agent" now means everything from a system prompt to a six-hour autonomous process. Here's a five-tier taxonomy based on autonomy, state, and blast radius — and why most agents on sale are tier one.

Most teams reach for fine-tuning when a decent retrieval pipeline and a tight prompt would have been cheaper, faster, and far easier to change tomorrow. Here's how to tell the difference.
Get started
Claim your address before someone else does — free to start, with an AI-native inbox built in.