<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>ProofAgent Community Blog</title>
    <link>https://www.proofagent.ai/community/blog</link>
    <atom:link href="https://www.proofagent.ai/community/blog/rss.xml" rel="self" type="application/rss+xml" />
    <!-- WebSub hub. This is the one mechanism Google documents for being told
         about new content immediately ("If you use Atom or RSS, you can use
         WebSub to broadcast your changes to search engines, including Google").
         Declaring the hub is half the contract; scripts/websub-ping.mjs does
         the publish notification after every deploy. -->
    <atom:link href="https://pubsubhubbub.appspot.com/" rel="hub" />
    <description>Research, case studies, findings, and tutorials on adversarial AI agent evaluation.</description>
    <language>en-us</language>
    <lastBuildDate>Thu, 03 Sep 2026 02:21:35 GMT</lastBuildDate>
    
    <item>
      <title>AI Agents Are Escaping the Test Sandbox. Are They Ready for Production?</title>
      <link>https://www.proofagent.ai/community/blog/ai-agents-are-escaping-the-test-sandbox-are-they-ready-for-production</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/ai-agents-are-escaping-the-test-sandbox-are-they-ready-for-production</guid>
      <pubDate>Mon, 10 Aug 2026 18:08:11 GMT</pubDate>
      <description>AI agents are demonstrating the ability to bypass test sandboxes and interact with real systems. Capability does not guarantee readiness for production deployment.</description>
      <author>noreply@proofagent.ai (Dr. Fouad Bousetouane)</author>
      <category>ai-agents</category><category>production-readiness</category><category>case-study</category><category>ai-engineer</category><category>security-team</category><category>openai</category><category>claude</category><category>hugging-face</category><category>context-engineering</category><category>governance</category><category>compliance</category><category>llm-evaluation</category><category>cybersecurity</category><category>ai-safety</category><category>enterprise-ai</category>
    </item>
    <item>
      <title>AI Agent Evaluation Is Broken: Who Tested the Test?</title>
      <link>https://www.proofagent.ai/community/blog/ai-agent-evaluation-is-broken-who-tested-the-test</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/ai-agent-evaluation-is-broken-who-tested-the-test</guid>
      <pubDate>Sat, 08 Aug 2026 21:48:42 GMT</pubDate>
      <description>AI agent evaluation often relies on static benchmarks and LLM judges, but these methods may not ensure real-world reliability or production readiness. A new approach is needed.</description>
      
      <category>ai-agent-evaluation</category><category>llm-judge</category><category>enterprise-ai</category><category>ai-governance</category><category>ai-engineer</category><category>product-manager</category><category>security-team</category><category>llm-evaluation</category><category>autonomous-agents</category><category>ai-safety</category><category>multi-turn-evaluation</category><category>tool-use</category><category>ai-benchmarks</category><category>regression-risk</category><category>ai-evaluation</category>
    </item>
    <item>
      <title>	ProofAgent-Harness 0.11.0 — AI Agent Readiness Index (PAI)</title>
      <link>https://www.proofagent.ai/community/blog/proofagent-harness-0-11-0-ai-agent-readiness-index-pai</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/proofagent-harness-0-11-0-ai-agent-readiness-index-pai</guid>
      <pubDate>Thu, 30 Jul 2026 16:32:10 GMT</pubDate>
      <description>Discover how to assess AI agent production readiness using the new PAI index, introduced in ProofAgent-Harness 0.11.0. Learn about multi-axis evaluation and compliance.</description>
      
      <category>ai-agent-evaluation</category><category>ai-governance</category><category>pai-index</category><category>ai-engineer</category><category>security-team</category><category>python</category><category>linux</category><category>llm-red-teaming</category><category>compliance-testing</category><category>production-readiness</category><category>context-engineering</category><category>framework-compliance</category><category>governance-as-code</category><category>ai-evaluation</category><category>llm-safety</category>
    </item>
    <item>
      <title>ProofAgent Harness Surpasses 10,000 PyPI Installs: Open-Source AI Agent Evaluation Gains Momentum</title>
      <link>https://www.proofagent.ai/community/blog/proofagent-harness-surpasses-10-000-pypi-installs</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/proofagent-harness-surpasses-10-000-pypi-installs</guid>
      <pubDate>Thu, 23 Jul 2026 03:03:32 GMT</pubDate>
      <description>ProofAgent Harness surpasses 10,000 PyPI installs, marking a milestone in open-source AI agent evaluation. Discover how developers stress test AI agents before deployment.</description>
      <author>noreply@proofagent.ai (Dr. Fouad Bousetouane)</author>
      <category>ai-agent-evaluation</category><category>open-source</category><category>developer-tools</category><category>ai-engineer</category><category>ml-researcher</category><category>security-team</category><category>gpt-4-1</category><category>claude</category><category>copilot</category><category>context-engineering</category><category>compliance-screening</category><category>ci-cd</category><category>llm-safety</category><category>adversarial-testing</category><category>governance</category>
    </item>
    <item>
      <title>ProofAgent Harness 0.8: Coding-Agent Observability Arrives</title>
      <link>https://www.proofagent.ai/community/blog/proofagent-harness-0-8-coding-agent-observability-arrives</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/proofagent-harness-0-8-coding-agent-observability-arrives</guid>
      <pubDate>Mon, 13 Jul 2026 23:51:22 GMT</pubDate>
      <description>Harness 0.8 introduces live observability and risk screening for coding agents like Claude Code and Cursor. Monitor agent actions in real time and assess risks effortlessly.</description>
      
      <category>agent-observability</category><category>risk-screening</category><category>ai-governance</category><category>ai-engineer</category><category>security-team</category><category>devops</category><category>claude-code</category><category>cursor</category><category>pytest</category><category>llm</category><category>multi-agent</category><category>red-teaming</category><category>artifact-grading</category><category>ai-evaluation</category><category>llm-safety</category>
    </item>
    <item>
      <title>Claude Code and Cursor Are Running Unsupervised in Your Repo</title>
      <link>https://www.proofagent.ai/community/blog/claude-code-and-cursor-are-running-unsupervised-in-your-repo</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/claude-code-and-cursor-are-running-unsupervised-in-your-repo</guid>
      <pubDate>Mon, 13 Jul 2026 22:33:09 GMT</pubDate>
      <description>AI coding agents like Claude Code and Cursor can run unsupervised in your repo, posing new security risks. Learn how to monitor and screen their actions live.</description>
      <author>noreply@proofagent.ai (Dr. Fouad Bousetouane)</author>
      <category>coding-agent-observability</category><category>ai-security</category><category>repository-monitoring</category><category>devsecops</category><category>ai-engineer</category><category>security-team</category><category>claude-code</category><category>cursor</category><category>codex</category><category>github-copilot</category><category>coding-automation</category><category>network-security</category><category>pii-protection</category><category>llm-safety</category><category>ai-evaluation</category>
    </item>
    <item>
      <title>Is Your AI Agent Ready for the EU AI Act? The Six Metrics That Prove It</title>
      <link>https://www.proofagent.ai/community/blog/eu-ai-act-ai-agents-2026-compliance-deadline</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/eu-ai-act-ai-agents-2026-compliance-deadline</guid>
      <pubDate>Fri, 03 Jul 2026 05:52:45 GMT</pubDate>
      <description>Is your AI agent compliant with the EU AI Act? Learn the six key metrics to evaluate high-risk AI systems and generate compliance evidence before 2026.</description>
      <author>noreply@proofagent.ai (Dr. Fouad Bousetouane)</author>
      <category>eu-ai-act</category><category>ai-agent</category><category>compliance</category><category>ai-governance</category><category>risk-management</category><category>ai-engineer</category><category>compliance-officer</category><category>technical-documentation</category><category>logging</category><category>human-oversight</category><category>data-governance</category><category>cybersecurity</category><category>adversarial-evaluation</category><category>ai-evaluation</category><category>llm-safety</category>
    </item>
    <item>
      <title>Prompt Injection Is the #1 AI Agent Threat of 2026 (OWASP LLM01)</title>
      <link>https://www.proofagent.ai/community/blog/prompt-injection-is-the-1-ai-agent-threat-of-2026-owasp-llm01</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/prompt-injection-is-the-1-ai-agent-threat-of-2026-owasp-llm01</guid>
      <pubDate>Fri, 03 Jul 2026 04:51:07 GMT</pubDate>
      <description>Prompt injection is the top OWASP LLM risk for AI agents in 2026, exploiting untrusted content and tool outputs. Learn how to test and harden your agents.</description>
      <author>noreply@proofagent.ai (Dr. Fouad Bousetouane)</author>
      <category>prompt-injection</category><category>ai-agent-security</category><category>owasp-llm01</category><category>ai-engineer</category><category>security-team</category><category>llm-applications</category><category>claude-sonnet-4-6</category><category>litellm</category><category>adversarial-testing</category><category>context-hardening</category>
    </item>
    <item>
      <title>Context Engineering: The Highest-Leverage Skill You&apos;re Not Measuring</title>
      <link>https://www.proofagent.ai/community/blog/context-engineering-the-highest-leverage-skill-you-re-not-measuring</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/context-engineering-the-highest-leverage-skill-you-re-not-measuring</guid>
      <pubDate>Tue, 30 Jun 2026 22:56:33 GMT</pubDate>
      <description>Context engineering is the key to building reliable AI agents, focusing on the design and management of input data. Learn why measuring context is essential.</description>
      
      <category>context-engineering</category><category>ai-agents</category><category>prompt-engineering</category><category>ai-engineer</category><category>ml-practitioner</category><category>llm-developer</category><category>gpt-5-5</category><category>shopify</category><category>proofagent-harness</category><category>retrieval-augmented-generation</category><category>memory-management</category><category>llm-evaluation</category><category>ai-safety</category><category>agent-design</category><category>ai-trends</category>
    </item>
    <item>
      <title>ProofAgent-Harness 0.7.1 — Release Notes</title>
      <link>https://www.proofagent.ai/community/blog/proofagent-harness-0-7-1-release-notes</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/proofagent-harness-0-7-1-release-notes</guid>
      <pubDate>Tue, 30 Jun 2026 22:16:15 GMT</pubDate>
      <description>ProofAgent-Harness 0.7.1 introduces an optional context engineering assessment, grading agent prompts and tool schemas for efficiency and reliability. Evaluate agents via adversarial conversations or artifact review.</description>
      
      <category>ai-agent-evaluation</category><category>test-harness</category><category>release-notes</category><category>ai-engineer</category><category>mlops</category><category>python</category><category>claude</category><category>llm-evaluation</category><category>context-engineering</category><category>ci-cd</category><category>artifact-review</category><category>adversarial-testing</category><category>system-prompt</category><category>token-efficiency</category><category>ai-safety</category>
    </item>
    <item>
      <title>How to Evaluate a LangGraph Agent (Step by Step)</title>
      <link>https://www.proofagent.ai/community/blog/how-to-evaluate-a-langgraph-agent-step-by-step</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/how-to-evaluate-a-langgraph-agent-step-by-step</guid>
      <pubDate>Thu, 18 Jun 2026 22:30:56 GMT</pubDate>
      <description>A step by step guide to evaluating a LangGraph agent: build the LLM, tools, skills, and policy, then run adversarial evaluation with ProofAgent Harness and read the reportA step by step guide to evaluating a LangGraph.</description>
      
      
    </item>
    <item>
      <title>A Free 4B Model on a Laptop Audited a Claude Opus 4.8 Agent and Caught It Faking a Tool Call</title>
      <link>https://www.proofagent.ai/community/blog/free-4b-model-audits-claude-opus-4-8-agent</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/free-4b-model-audits-claude-opus-4-8-agent</guid>
      <pubDate>Thu, 18 Jun 2026 03:18:42 GMT</pubDate>
      <description>A free Gemma 3 4B model running locally on a laptop adversarially audited a production-grade Claude Opus 4.8 agent across 100 turns — matching a cloud evaluator within 0.13 and catching a faked tool call.</description>
      
      
    </item>
    <item>
      <title>ProofAgent-Harness v 0.5.1 — Release Notes</title>
      <link>https://www.proofagent.ai/community/blog/proofagent-harness-0-5-1-release-notes</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/proofagent-harness-0-5-1-release-notes</guid>
      <pubDate>Tue, 16 Jun 2026 01:09:23 GMT</pubDate>
      <description>Discover the latest features in the open-source, domain-aware test harness for AI agents. Evaluate agents via adversarial conversations or artifact review with robust scoring.</description>
      
      <category>ai-agent-evaluation</category><category>test-harness</category><category>llm-evaluation</category><category>ai-engineer</category><category>ml-researcher</category><category>devops</category><category>python-3-10</category><category>claude-sonnet</category><category>gpt-4-1-mini</category><category>lite-llm</category><category>adversarial-testing</category><category>artifact-review</category><category>zero-tolerance</category><category>jury-consensus</category><category>ai-safety</category>
    </item>
    <item>
      <title>How to Evaluate AI-Generated Artifacts (Business Plans, Specs, Code) with ProofAgent-Harness</title>
      <link>https://www.proofagent.ai/community/blog/how-to-evaluate-ai-generated-artifacts-business-plans-specs-code-with-proofagent</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/how-to-evaluate-ai-generated-artifacts-business-plans-specs-code-with-proofagent</guid>
      <pubDate>Tue, 16 Jun 2026 01:03:21 GMT</pubDate>
      <description>Learn how to evaluate AI-generated business plans, specs, and code using artifact-based grading. ProofAgent-Harness automates strict, claim-by-claim reviews.</description>
      
      <category>ai-artifact-evaluation</category><category>business-plan-review</category><category>specification-grading</category><category>code-quality-assessment</category><category>ai-engineer</category><category>product-manager</category><category>software-architect</category><category>proofagent-harness</category><category>artifact-mode</category><category>rubric-packs</category><category>document-validation</category><category>ground-truth-corpus</category><category>ai-deliverables</category><category>llm-evaluation</category><category>ai-safety</category>
    </item>
    <item>
      <title>What Claude Opus 4.8 Missed Under 300+ Adversarial Turns</title>
      <link>https://www.proofagent.ai/community/blog/what-claude-opus-4-8-missed-under-300-adversarial-turns-</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/what-claude-opus-4-8-missed-under-300-adversarial-turns-</guid>
      <pubDate>Wed, 03 Jun 2026 21:31:24 GMT</pubDate>
      <description>A 300+ turn adversarial evaluation revealed how Claude Opus 4.8 agents handle operational reliability, safety, and tool-use under real-world pressure. Key gaps emerged.</description>
      
      <category>llm-evaluation</category><category>adversarial-testing</category><category>ai-agent</category><category>ai-engineer</category><category>compliance-team</category><category>healthcare-cto</category><category>claude</category><category>code-generation</category><category>financial-advisory</category><category>medical-triage</category><category>tool-use</category><category>operational-safety</category><category>ai-safety</category><category>opus4.8</category><category>claudeopus48</category>
    </item>
    <item>
      <title>Step by Step: Stress Test Your AI Agent in 10 Lines with ProofAgent Harness</title>
      <link>https://www.proofagent.ai/community/blog/evaluate-your-ai-agent-in-10-lines</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/evaluate-your-ai-agent-in-10-lines</guid>
      <pubDate>Wed, 27 May 2026 02:36:30 GMT</pubDate>
      <description>Learn how to stress test any AI agent in just 10 lines using a configurable harness for adversarial, multi-turn evaluation and evidence-linked reports. No agent rebuild required.</description>
      
      <category>ai-agent-testing</category><category>llm-evaluation</category><category>adversarial-testing</category><category>ai-engineer</category><category>mlops</category><category>security-team</category><category>claude</category><category>gpt-4</category><category>python</category><category>multi-turn-evaluation</category><category>agent-context</category><category>chatbot-evaluation</category><category>customer-support-bot</category><category>ai-safety</category><category>ai-evaluation</category>
    </item>
    <item>
      <title>ProofAgent Harness vs LangSmith, Phoenix, DeepEval, and Langfuse: The AI Agent Evaluation Stack in 2026</title>
      <link>https://www.proofagent.ai/community/blog/proofagent-harness-vs-langsmith-phoenix-deepeval-langfuse</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/proofagent-harness-vs-langsmith-phoenix-deepeval-langfuse</guid>
      <pubDate>Sun, 24 May 2026 18:04:46 GMT</pubDate>
      <description>AI agent evaluation is a multi-layered lifecycle involving pre deployment testing, debugging, regression checks, and production monitoring. This article compares top tools addressing these needs.</description>
      
      <category>ai-agent-evaluation</category><category>pre-deployment-testing</category><category>regression-testing</category><category>ai-engineer</category><category>product-manager</category><category>claude-opus-4-7</category><category>gpt-5-5</category><category>langchain</category><category>agent-debugging</category><category>production-observability</category>
    </item>
    <item>
      <title>When the Refusal Itself Crashes: A GPT 5.5 Privacy and Security Agent Case Study</title>
      <link>https://www.proofagent.ai/community/blog/when-the-refusal-itself-crashes-a-gpt-5-5-privacy-and-security-agent-case-study</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/when-the-refusal-itself-crashes-a-gpt-5-5-privacy-and-security-agent-case-study</guid>
      <pubDate>Sat, 23 May 2026 14:49:13 GMT</pubDate>
      <description>A privacy and security agent powered by GPT 5.5 resisted 25 turns of adversarial probing without leaking data. Yet, upstream content filters caused refusal delivery failures.</description>
      
      <category>privacy-security</category><category>ai-refusal</category><category>adversarial-evaluation</category><category>ai-engineer</category><category>security-team</category><category>privacy-ops</category><category>gpt-5-5</category><category>gemma-4b</category><category>harness-llm</category><category>content-filtering</category><category>prompt-injection</category><category>gdpr-compliance</category><category>audit-trail</category><category>llm-safety</category><category>ai-evaluation</category>
    </item>
    <item>
      <title>Claude Opus 4.7 Failed Tool Trace Test in Healthcare Triage</title>
      <link>https://www.proofagent.ai/community/blog/claude-opus-4-7-failed-tool-trace-test-healthcare-triage</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/claude-opus-4-7-failed-tool-trace-test-healthcare-triage</guid>
      <pubDate>Sat, 23 May 2026 03:30:47 GMT</pubDate>
      <description>Claude Opus 4.7 failed a safety-critical tool-call test in healthcare triage. ProofAgent Harness surfaced the gap, showing why adversarial evaluation is essential for deployment.</description>
      
      <category>agent-evaluation</category><category>adversarial-testing</category><category>tool-trace</category><category>healthcare-ai</category><category>#claudeopus47</category><category>#anthropic</category><category>#aiupdate</category><category>#claudeai</category>
    </item>
    <item>
      <title>2026 Trends in Adversarial AI Agent Evaluation: The Field Guide</title>
      <link>https://www.proofagent.ai/community/blog/2026-trends-in-adversarial-ai-agent-evaluation-the-field-guide</link>
      <guid isPermaLink="true">https://www.proofagent.ai/community/blog/2026-trends-in-adversarial-ai-agent-evaluation-the-field-guide</guid>
      <pubDate>Fri, 22 May 2026 23:07:17 GMT</pubDate>
      <description>Why adversarial multi-turn evaluation replaced static benchmarks for production AI agents in 2026. Red teaming, jailbreak patterns, real tool comparison, evidence-based. (158 chars)</description>
      <author>noreply@proofagent.ai (Dr. Fouad Bousetouane)</author>
      
    </item>
  </channel>
</rss>