Read the source, then write down what it actually says
I build the evidence layer for AI agents, and I write about where it does not
hold yet. Every figure below was read at its primary source, every note lists what it was
checked against, and where a source did not support a claim the claim is not here.
The App Defense Alliance AI Agent Specification defines three kinds of lab evidence and says an untagged requirement defaults to black-box testing. None of its 37 requirements is tagged, so two that need a configuration file or system logs read as black-box passes. Filed as issue 514.
The MCP TypeScript SDK's 1.30.1 patch added a 4 MiB body limit and a 100 message batch limit. On the Express path the SDK documents, the batch limit holds and the body limit is never consulted. A 100 kilobyte parser default refuses first, and the transport never hears about it.
Proof-of-Control positions itself as the evidence half of agent assurance, not runtime enforcement. Its own threshold requires an in-path gateway that withholds actions, and its threat model has no row for that gateway. Filed as issue 78.
Docker published two more CVEs against its agent sandbox on 15 September. Read alongside the three it published in June and August, all five sit in the two controls its README sells, and three of them fail the same way.
I ran the MCP Inspector's skill verifier against a server I wrote. A skill that advertises no digests gets the same word, the same JSON and the same exit code as one whose every file hashed clean.
Four hours apart, awslabs/mcp merged a parser-based read-only policy for one database server and twenty more regex keywords for another. I installed the published package and found the keyword answer still has a gap the parser answer cannot have.
DDRop landed on 14 September and walks past the platform appraisal policy I shipped three weeks ago. Reading it sent me to a paper from last year that had already walked past the same check, by the same four people, for under $50.
CVE-2026-86121 is filed as missing authentication in a computer-use agent sandbox, critical, fixed in 0.3.42. The release named as the fix changes which interface the server listens on. Its authentication path is byte-identical to the version before it, and the same allow-all branch shipped again yesterday.
MCP 2026-07-28 made it a protocol violation for a server to change its tool list as a side effect of what you just did. The same release forbade the server from telling you the list changed unless you subscribed, and neither official SDK subscribes by default.
One coding agent answered repo-controlled git config with a confirmation prompt in August. Another answered the neighbouring case with a refusal. The gap between those two choices is the whole argument, and neither public record tells you which one you are running.
A gap in the SEAT post-handshake attestation draft turned out to have been argued on the working group list in July, and specified in a closed issue a year before that. What I filed instead came from reading the RFCs the draft already cites.
The MCP tasks extension requires an authorization check on every task request. Its own rationale says that check is often impossible, the sentence that used to make servers disclose that is gone, and no error code represents a denial. One of the two losses was already found and fixed by somebody else three weeks ago.
Spring AI released a fix on 21 August for a bug where a tool absent from the request could still be called. The vulnerable dispatch is one line. The public fix exists on one of the three affected release lines, and the fallback left almost nothing useful in the log.
Publishing proves you own the namespace and you own the package. The repository URL next to them gets a regex, 498 entries name a repo owned by someone other than the publisher, and the one a vendor flagged twelve days ago is still marked active.
Google Cloud tells you to update go-sev-guest so your parser stops breaking on v4 attestation reports. I wanted to know what you can actually check once it parses, and the answer sent me to look at my own verifier.
Thirty poisoned npm packages are gone along with their attestations, and the Sigstore copy that survives can only be fetched with a digest that no longer exists in any public place I could find.
Coverage of the reasoning-trace harvest reported three different totals as if they disagreed. They are all the paper's own numbers, at three different denominators.
The gym booking story is being told as an API with no authorization. The agent's own account says two of the three operations it touched returned 403, and one did not.
agent-securityapiauthorization
Essays
Longer pieces, published on dev.to and cross-posted to Medium.