Subscribe to get new posts from mleg.tech in your reader of choice.
Your agent hallucinates a contract clause. Sends it to a customer. Then, an hour later, a dashboard somewhere turns red.
Uptime and latency can't tell you if an agent's output is correct. Why AI observability needs a different kind of tooling than the APM stack you already have.
Bedrock Agents is now Classic, closed to new customers, and in maintenance mode — what AgentCore actually changes, and what it means if you're already running agents in production.
Reducer bugs, state bloat, and the deep-merge fix that took a full day to trace — lessons from running multi-agent workflows on AWS Bedrock at scale.
Why fixed-size chunking quietly tanks retrieval quality, and what content-hash dedup buys you at scale.
How an eval pipeline caught stale-context regressions before they reached production, and what it cost to build.
<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom"><title>mleg.tech — Michael Legemah</title><subtitle>Notes on building AI systems that ship.</subtitle><link href="https://mleg.tech/feed.xml" rel="self"/><link href="https://mleg.tech/"/><id>https://mleg.tech/</id><updated>2026-08-25T09:00:00.000Z</updated><author><name>Michael Legemah</name></author><entry><title>Your eval dashboard is a crime scene photo</title><link href="https://mleg.tech/blog/your-eval-dashboard-is-a-crime-scene-photo"/><id>https://mleg.tech/blog/your-eval-dashboard-is-a-crime-scene-photo</id><published>2026-08-25T09:00:00.000Z</published><updated>2026-08-25T09:00:00.000Z</updated><summary>Your agent hallucinates a contract clause. Sends it to a customer. Then, an hour later, a dashboard somewhere turns red.</summary><category term="MCP"/></entry><!-- 5 more entries --></feed>