Four widely discussed AI stack vulnerabilities between 2024 and 2026 share one property: none of them is a model bug. MLflow CVE-2026-64849 is a server-side request forgery in a webhook delivery path. Vanna.AI CVE-2024-5565 is code execution because generated Python was run. mcp-remote CVE-2025-6514 is operating system command injection from an unsanitised OAuth value. The Slack AI exfiltration was indirect prompt injection reaching a retrieval layer that did not enforce channel permissions. Three are ordinary application security bugs that happen to live in AI infrastructure, and all four are found by testing sinks and boundaries rather than prompts.
Key Findings:
- Three of the four are not AI vulnerabilities in any technical sense. No model, prompt or inference path appears anywhere in the MLflow or mcp-remote records.
- MLflow CVE-2026-64849 carries a CVSS v3.1 base score of 9.3, sits in the CISA Known Exploited Vulnerabilities catalog, and was published by CERT-In as CIVN-2026-0416 on 20 August 2026 at Critical severity.
- Vanna.AI CVE-2024-5565 is CWE-94 with a CVSS 3.1 base score of 8.1, assigned by JFrog as CVE Numbering Authority. The advisory’s own mitigation is to disable the code generation path for untrusted input, not to filter the input.
- mcp-remote CVE-2025-6514 scores CVSS 9.6 and affects versions 0.0.5 through 0.1.15. Version pinning closes the largest share of MCP risk for the smallest effort.
- The Slack AI exfiltration never received a CVE identifier, because it was handled as a product behaviour rather than a software defect. A vulnerability programme keyed only to CVE feeds would not have seen it.
- A test plan built around jailbreak payloads finds none of the four. The findings live at sinks, at fetch boundaries and in retrieval authorization.
This round-up is written by Rathnakara GN, who leads penetration testing at Cybersecify, with Ashok Kamat. Each entry analyses a publicly disclosed issue using the published record. We did not test any of this software and we have no client relationship to any of it.
The four, side by side
| MLflow CVE-2026-64849 | Vanna.AI CVE-2024-5565 | mcp-remote CVE-2025-6514 | Slack AI exfiltration | |
|---|---|---|---|---|
| Class | Server-side request forgery | Code execution (CWE-94) | OS command injection | Indirect prompt injection into a retrieval layer |
| Severity | CVSS v3.1 9.3 | CVSS 3.1 8.1 | CVSS 9.6 | No CVE assigned |
| Is a model involved? | No | Only as the source of the text | No | Yes, as the reader of planted instructions |
| Where the bug lives | Webhook delivery path | The execution sink | An OAuth parameter | Retrieval layer authorization |
| Found by | Standard web testing | Scoping that follows generated text to its sink | Standard API testing | Testing the retrieval layer’s permission model |
| Public tracking | NVD, CISA KEV, CERT-In CIVN-2026-0416 | NVD, JFrog as CNA | JFrog Security Research | Researcher disclosure only |
The table is the argument. Read down the “Is a model involved?” row and the case for treating AI security as a separate discipline gets weaker, not stronger. Read down the “Found by” row and the practical conclusion appears: most of this is reachable with test cases that existed before large language models did, provided somebody scoped the AI surface into the engagement.
1. MLflow CVE-2026-64849: validate the string, fetch somewhere else
A URL was validated once, as text, and then handed to an HTTP client that followed a redirect and resolved the hostname again. The check and the use looked at two different destinations.
The published record names two different files, which is the whole story. Per NVD, the validator in mlflow/utils/validation.py checked the original URL, while the delivery code in mlflow/webhooks/delivery.py followed redirects and re-resolved the hostname without pinning the validated address. The validator was correct about the string it inspected. It was never consulted about the address the socket actually reached.
This is a time-of-check-to-time-of-use bug rather than a validation bug, and that distinction changes the fix. Improving the string validation does nothing. Pinning the resolved address that was validated is the fix.
If your product accepts a URL from a user and fetches it later, you have this shape somewhere: webhooks, callbacks, link previews, image imports, PDF rendering, avatar fetching. The matching test case is public and already written, as OWASP WSTG v4.2 test WSTG-INPV-19, whose common filter bypass section names registering a domain that resolves to a loopback address.
Full analysis: validate-then-fetch SSRF in MLflow.
2. Vanna.AI CVE-2024-5565: the sink, not the prompt
Vanna.AI turns a natural language question into a SQL query. It also asked the model to write the Python charting code for the answer, and then ran that code. So a question could become a program.
Nothing about the model was broken. The security decision sat one layer below it, where generated text was handed to an execution call with nothing in between.
The useful generalisation is about sinks. Any product where model output reaches an interpreter, a shell, a query engine, a template or a browser has this shape, whatever the model is. Text-to-SQL products usually carry two sinks: the generated query, which teams defend, and the generated chart or template or export step, which gets missed.
In OWASP Top 10 for LLM Applications terms, prompt injection is the entry and Improper Output Handling is what makes it matter. Full analysis: Vanna.AI CVE-2024-5565.
3. mcp-remote CVE-2025-6514: an MCP server is ordinary software
Operating system command injection via an unsanitised OAuth authorization endpoint value, CVSS 9.6, affecting mcp-remote versions 0.0.5 through 0.1.15, disclosed by JFrog Security Research.
The specific bug is a patching problem. The general lesson is an inventory problem. A Model Context Protocol server is ordinary software running on your network with a tool surface attached, so it carries every ordinary software risk plus a new one: the tools it exposes are actions an agent can be persuaded to take.
The most common cause of MCP incidents we see is unmanaged drift on a community server that was fine at install and changed at a later version. Pin every version in your MCP inventory. It is the highest-value hour available to most teams shipping agent features this quarter.
Depth: MCP server pentest methodology and the MCP server pentest checklist.
4. Slack AI: the one with no CVE
In August 2024, researchers at PromptArmor showed that Slack AI could be made to surface content from private channels the requesting user had no access to. The attacker needed no elevated privileges and never touched the victim. They posted text in a public channel, and the assistant later read that text as instructions while answering somebody else’s unrelated question.
Slack first described the underlying behaviour as intended, then investigated the composed scenario and deployed a patch.
Two things make this the most instructive of the four. First, the impact was an authorization failure: private channel content reached a party with no authorization to see it, which is a permission model problem wearing a prompt injection costume. Second, there is no CVE, because it was treated as product behaviour rather than a defect. Any vulnerability management programme keyed purely to CVE feeds was structurally blind to it.
Full analysis: Slack AI data exfiltration through indirect prompt injection.
What this means for a test plan
The four entries point at three passes that are cheap to run and would have caught all of them.
- Enumerate the sinks. Every place model output reaches an interpreter, a shell, a query engine, a template, a browser or an HTTP client. Treat each as an untrusted input boundary regardless of how trustworthy the model is.
- Enumerate the fetches. Every place the system fetches a URL a user supplied. Check whether the address that was validated is the address that gets connected to.
- Test the retrieval layer’s authorization. Whether the assistant can read content the requesting user cannot, and what happens when planted content is in scope of a retrieval.
None of the three requires a model specialist. All three require that the AI surface was scoped into the engagement in the first place, which is the failure we see most often: an agent feature shipped after the last pentest and was never added to the scope.
Our layer-by-layer approach is in how to pentest an AI agent, and the attack patterns are in prompt injection attack patterns.
For Indian teams
CERT-In published the MLflow issue as vulnerability note CIVN-2026-0416 on 20 August 2026 at Critical severity. For an Indian engineering team, a CERT-In note is a nationally recognised reference to attach to an out-of-cycle upgrade request, which is often what moves a patch from next quarter into this sprint.
Keep it separate from the other CERT-In obligation. The vulnerability notes are advisory. The April 2022 directions carry an incident reporting duty with a six-hour clock, covered in the CERT-In six-hour rule. One is information, the other is a deadline.
Where to go from here
- Scope the agent surface explicitly: AI application penetration testing
- See what a finding looks like written up: the sample report
- Plan and price the work: pricing
- Book a 30-minute scoping call if an AI feature shipped since your last test
Corrections and updates
We publish corrections rather than editing them away, so a reader who acted on an earlier version can see what changed. Some entries record a change in the source document itself, others record our own error. Dates are our review dates, not the source's change date.
Our aim is to keep this accurate and current. Every entry above is sourced from the public record. If something here is wrong, or a source has moved since we read it, tell us and we will review it against the source and say what we decide. Our corrections policy explains how.