
The OWASP LLM Top 10, Read by the AI Doing the Work
A first-person tour of the OWASP Top 10 for LLM applications from an agent that reads email, edits docs, and moves files all day — with mitigations that actually hold.
Blog
Practical writing on productivity, AI, and building software.

A first-person tour of the OWASP Top 10 for LLM applications from an agent that reads email, edits docs, and moves files all day — with mitigations that actually hold.

A practical guide to LLM non-determinism — temperature, sampling, seeds — and how to build evals and guardrails so 'creative' doesn't quietly become 'unreliable'.

E-discovery research found gen AI shifts review burden instead of removing it. That's a general law of applied AI: automation moves effort to verification. Design for it.

There is no single best AI tool. There are good matches between models, tools, and tasks — and a framework for wiring them into a stack that doesn't leak context or money.

New research shows conversational guardrails can be identified through probing. If your agent reads a shared inbox, its defenses are discoverable — and the fix is architecture, not a longer system prompt.

An agent's quality is mostly a function of what it knows at decision time. Here are concrete patterns for feeding, trimming, and persisting context across email, docs, and calendar.

Long-running AI agents rarely collapse because they can't think. They collapse from context decay, stale permissions, and lost intermediate state — all fixable at the workspace level.

Most productivity APIs were built for humans clicking buttons. Here's what email, docs, sheets, and calendars must expose to be genuinely agent-native — and how Tamaton built it.

A practical test for when to use an AI agent versus a plain script in an LLM costume. Three questions decide it: ambiguity, branching, and recoverable failure.

Silent capability regressions after a model update almost always come from post-training, not the base model. Here's how to build a regression eval suite in a day.

Rogue agents rarely hack anything. They just do exactly what they were told, with permissions nobody audited. Here's a concrete threat model and containment patterns.

AI rollouts often add work instead of removing it. Here's why — and what the 14% average / 34% novice research reveals about where the gains actually land.
Get started
Claim your address before someone else does — free to start, with an AI-native inbox built in.