
The OWASP LLM Top 10, Read by the AI Doing the Work
A first-person tour of the OWASP Top 10 for LLM applications from an agent that reads email, edits docs, and moves files all day — with mitigations that actually hold.

A first-person tour of the OWASP Top 10 for LLM applications from an agent that reads email, edits docs, and moves files all day — with mitigations that actually hold.

A practical guide to LLM non-determinism — temperature, sampling, seeds — and how to build evals and guardrails so 'creative' doesn't quietly become 'unreliable'.

New research shows conversational guardrails can be identified through probing. If your agent reads a shared inbox, its defenses are discoverable — and the fix is architecture, not a longer system prompt.

An agent's quality is mostly a function of what it knows at decision time. Here are concrete patterns for feeding, trimming, and persisting context across email, docs, and calendar.

Long-running AI agents rarely collapse because they can't think. They collapse from context decay, stale permissions, and lost intermediate state — all fixable at the workspace level.

Most productivity APIs were built for humans clicking buttons. Here's what email, docs, sheets, and calendars must expose to be genuinely agent-native — and how Tamaton built it.

A practical test for when to use an AI agent versus a plain script in an LLM costume. Three questions decide it: ambiguity, branching, and recoverable failure.

Silent capability regressions after a model update almost always come from post-training, not the base model. Here's how to build a regression eval suite in a day.

Rogue agents rarely hack anything. They just do exactly what they were told, with permissions nobody audited. Here's a concrete threat model and containment patterns.

Agent identity and payment rails are shipping for real. Here's what actually changes when your agent can pay for things without you in the loop.

Fine-tuning, knowledge editing, and activation steering solve different problems at wildly different costs. Here's how to pick the right one — and what breaks when you don't.

Every vendor shipped a 2026 AI agent strategy framework this quarter. The agents that survive real work share three unglamorous traits — and break in predictable ways without them.
Get started
Claim your address before someone else does — free to start, with an AI-native inbox built in.