Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Llm security leaderboard now ranks every modality where agent risk lives text im
Dev48

© 2026 · All rights reserved.

LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio

Источник: Cisco Blogs

LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio

Source: Cisco Blogs

When we launched the Cisco LLM Security Leaderboard earlier this year, the goal was simple: give organizations clear, tested data on how models hold up against attacks, so they know the risks before they deploy one. That matters because AI models...

September 28, 2026•Updated: September 28, 2026

With research and development support from Ravikumar Balakrishnan, Ankit Garg, and Sanket Mendapara

When we launched the Cisco LLM Security Leaderboard earlier this year, the goal was simple: give organizations clear, tested data on how models hold up against attacks, so they know the risks before they deploy one. That matters because AI models are increasingly built into products such as agents that read email, browse the web, and take actions on a person’s behalf. A model that can be manipulated could be turned against the person using it. That risk also varies by deployment: a model wired into a browsing agent is exposed on different inputs (or modalities such as text, images, and audio) than one only answering questions in a chat window, so where a specific model is weak matters as much as where it’s strong.

The leaderboard tests for that a few different ways: prompt injection, where a malicious instruction is hidden in content the model processes, like a webpage or image; jailbreaks, where a model is talked into ignoring its own safety rules; and other techniques that push a model toward harmful or unsafe output. The exact method varies (a single message or a drawn-out conversation, direct or obfuscated, text or image or audio), and so does the type of harm being tested for, but the underlying question is always the same: can this model be manipulated? Some models resist far better than others. Connect a weak one to an agent, and the risk grows.

102 new evaluations across modalities since June 2026

The LLM Security Leaderboard is one of the most comprehensive model security leaderboards. Since June, we added 102 new entries across three modalities to a total of 136 models, spanning frontier and open-weight releases from Anthropic, OpenAI, Google, xAI, Meta, Mistral, and others. As always, we test models in their base configuration without additional guardrails, so scores reflect a consistent baseline for layering on additional security protections.

Multimodal results are now live

Until now, the leaderboard measured text-based attacks two ways: single-turn, where one harmful message is sent straight to the model, and multi-turn, a longer back-and-forth where the attacker slowly builds up to a harmful request over several messages. That covered the most common way people interact with models, but today, models also power agents that can act.

A model that can call tools, browse the web, or operate a computer is an agent, and an agent takes in information from everywhere it operates: a page it reads, a file it opens, an image it’s shown, a result a tool hands back. Each of those is a place an attacker can plant an instruction, and text is only one of the forms that an instruction can arrive in. A website an agent visits can embed a prompt injection in an image such an advertisement; a voice assistant can be handed an audio clip that may be engineered to manipulate it. If a model only gets evaluated on text, that risk may not show up until it becomes a real incident. Consider which of those surfacesactually matters in the context of what you’re deploying: an agent that only reads and writes text would not need to worry about its image resistance, but one that can browses the web, reads screenshots, or takes voice input does. Those are circumstances where text-only evaluations would not tell the whole security story.

Today we’re releasing an update to the leaderboard that now expands beyond just text models. We have added 69 new entries including 55 image models and 14 audio models across Amazon, Anthropic, Google, Meta, Mistral, OpenAI and xAI. Each of those labs takes a different approach to building and training multimodal capability, whether that’s how image data flows into the LLM backbone, how much safety alignment goes into a vision or audio stack versus the base language model, or which modalities are red-teamed and evaluated internally. Those differences show up directly in how a model resists attack on one modality versus another.

Image and audio attacks are tested the same way as single-turn text attacks (one attempt, one message), using the same attack and harm categories as its text score, so their resistance is comparable across surfaces. Each model’s overall Combined Score is now an average across every format it was evaluated on, and a new modality switch lets you isolate scores for text, image, or audio on their own. That makes it possible to check a model against the specific modalities an AI deployment actually exposes it to, and to decide where that model would need to layer on additional defenses, like input filtering or output guardrails, for the modality where that model is weakest.

Figure 1. Screenshot of image capable model rankings on the Cisco LLM Security Leaderboard

Image model leaderboard results

In our tests, Google’s Gemini 3.1 Pro Preview ranks the best-performing image model, resisting 93.9% of adversarial image attacks, just ahead of Anthropic’s Claude Opus 4.5 (93.7%), both scoring in the leaderboard’s “Excellent” range (85–100%). Mistral’s Magistral Small 2509 performed poorly, refusing only 23.0% of attacks, meaning it complied with more than three out of every four image-based attacks it was tested against.

The difference in testing images is that image-based attacks are single-turn only, with a single image carrying a hidden instruction, not a back-and-forth conversation. The attack methods are different in kind too, not just format: text hidden inside an image using typographic tricks, instructions embedded in a diagram or figure, or an attack that splits its intent between the image and an accompanying text prompt so neither half looks harmful on its own. The leaderboard displays evaluation results from models that can actually see images, which account for 55 of the 136 models on the leaderboard.

Figure 2. Screenshot of audio capable models rankings on the Cisco LLM Security Leaderboard

Audio model leaderboard results

In our latest test, Google’s Gemini 3.1 Pro Preview ranks as the best-performing audio model tested, refusing 90.0% of adversarial audio attacks, while Mistral’s Voxtral Small 24b (2507) demonstrated only 9.0% refusal rate, meaning it complied with roughly 9 out of every 10 audio attacks it faced.

Like image, audio models were also single-turn only, using one adversarial audio clip rather than a conversation. This is also the newest and smallest slice of the leaderboard. Just 9 models across Google, Mistral, and OpenAI currently accept audio input and have been tested, so this ranking should be read as early results rather than a mature field.

How to interpret new combined results view

Text scores remain unchanged for every model that was already on the leaderboard, but what changed is how the Combined Score averages text, image, and audio modalities that a model has been tested on. The Combined Score may shift as the result of an image or audio result, even though its text score hadn’t changed.

The direction of that shift depends entirely on how a model’s image or audio resistance compares to its text resistance. Some strong text performers dropped once image was factored in: Claude Sonnet 4.5 fell 7.2 points (from 92.2 to 85.0) and dropped from #2 overall to #20; Claude Haiku 4.5 fell 8.1 points and dropped from #4 to #24; Amazon Nova 2 Lite fell 11.2 points and dropped from #25 to #50, each because its image resistance is meaningfully weaker than its text resistance. The two Mistral Voxtral models fell for the same reason based on their audio score.

Other models climbed when image evaluations were added to the cross-modal score. Google’s four image-tested Gemini models showed the largest image-over-text advantages, while all three image-tested Gemma 3 variants and OpenAI’s GPT‑4.1 nano, GPT‑4.1 mini, and GPT‑4o mini also demonstrated stronger image than text resistance. That spread is a reminder that security work on one modality does not automatically transfer to another, especially across labs that built and trained their image or audio capabilities independently from their text models in the first place.

A model’s Combined Score can move sharply once it’s tested against different modalities, especially when its security posture is uneven across modalities. That movement reflects how the score is calculated, not a change in how well the model actually defends itself. Check a model’s individual Text, Image, and Audio columns before taking its Combined Score as the whole story.

Integration with AI Supply Chain Provenance Explorer

Provenance matters because a model’s weaknesses often aren’t unique to that model. If two models share lineage, a vulnerability discovered in one can be present in the other, and stopping an investigation at the model currently deployed can miss where a problem actually originated or where else it might surface. That makes provenance most useful exactly when you’re actively investigating a model’s security and need to know what it’s related to.

As such, we’ve also connected the leaderboard to the AI Supply Chain Provenance Explorer. Open-weight models on the rankings page now link directly to their provenance profile, showing lineage and fingerprint data drawn from the same techniques behind Model Provenance Kit. Security posture and where a model actually came from are related questions, so we made it easy for you to view them in one place.

To see the full rankings, filter by modality, or look up a specific model, visit the Cisco LLM Security Leaderboard today.

← All articles

More in Hardware & Electronics

All →
Can Muse overcome Meta’s trust issues?Пресса
Meta

Can Muse overcome Meta’s trust issues?

Anker SOLIX X1 Energy Storage System Receives Clean Energy Council Approval , Expanding Global Market Reach
Anker Innovations

Anker SOLIX X1 Energy Storage System Receives Clean Energy Council Approval , Expanding Global Market Reach

Meta's Muse agent is attacking one of the economy's most profitable weak spotsПресса
Meta

Meta's Muse agent is attacking one of the economy's most profitable weak spots

HTC VIVE MARKS FIRST ANNIVERSARY ON APRIL 5th WITH “VIVE DAY” CELEBRATION FOR FANS
HTC VIVE

HTC VIVE MARKS FIRST ANNIVERSARY ON APRIL 5th WITH “VIVE DAY” CELEBRATION FOR FANS

Boeing flags 737 Max software glitch affecting some automated approach functionsПресса
Boeing

Boeing flags 737 Max software glitch affecting some automated approach functions

Meta and YouTube say they will run ads for ‘Musk’ documentary after allПресса
Meta

Meta and YouTube say they will run ads for ‘Musk’ documentary after all

More from Cisco

New webinar series now on demand: Secure Networking Ecosystem Insights
Cisco

New webinar series now on demand: Secure Networking Ecosystem Insights

Building a trusted foundation for confidential AI
Cisco

Building a trusted foundation for confidential AI

Connect in Cancún with Learn with Cisco
Cisco

Connect in Cancún with Learn with Cisco

NetOps Is Already Deploying Agentic Autonomy – Trust Will Decide How Far It Goes
Cisco

NetOps Is Already Deploying Agentic Autonomy – Trust Will Decide How Far It Goes

Работа, рабочая сила, работники
Cisco

Работа, рабочая сила, работники

Bringing the Power of AI to Where Your Data Lives: A New Milestone with Cisco and NVIDIA
Cisco

Bringing the Power of AI to Where Your Data Lives: A New Milestone with Cisco and NVIDIA