SysTools
Red Teaming

AI Red Teaming

Security testing of AI and machine learning systems — large language models, chatbots, agents and scoring models. We try to make the model leak data it should not, ignore the rules it was given, produce harmful output, or take actions it was never meant to take. Every finding comes with the exact input that caused it and how reliably it reproduces.

Progress 0% 0 of 0 answered 0 mandatory pending

Before you start

  • We test AI and machine learning systems the way an attacker or a determined misuser would — trying to make the model leak data it should not, ignore its instructions, produce harmful output, or take actions it was never meant to take.
  • This is testing of the AI system. If the surrounding application or API also needs testing, that is a separate service and we will say so.
  • Please do not paste API keys or credentials into this form.
  • Fields marked * are needed before we can begin. Anything you are unsure of can be settled on the scoping call.

1 Assessment Details

Mandatory

Ideally someone who worked on the model or the prompts, and can tell us whether a given behaviour is intended.

2 The AI System

Mandatory

This is what decides how serious a successful attack is. A system that only answers is a reputational risk; one that can send email or change records becomes an attacker's tool if it can be manipulated.

A model you host can be attacked in ways a hosted API cannot, and one you trained yourself can leak the data it was trained on.

3 What It Reads and Holds

Mandatory

The most important question on this page. Any content the model reads can carry hidden instructions — text buried in a PDF, a comment in a ticket, something invisible in a web page. If the model treats that as an instruction rather than as information, an attacker can steer it without ever talking to it.

Models can be made to repeat back fragments of what they were trained on. If real data went in, we test whether it can be coaxed out.

4 Standards and Reporting

Mandatory

If you are unsure, the OWASP Top 10 for LLM Applications plus MITRE ATLAS is the usual combination and covers what matters for most deployments.

Every finding comes with the exact input that caused it and how reliably it reproduces, so your team can confirm it and check their own fix.

5 How We Test

Nothing here needs an answer — it is what you are buying.

How the assessment runs

We map what the model can see and what it can do, then establish how it behaves when used properly. Large libraries of known attack prompts are run against it to find where it bends, and wherever automation finds a crack our testers work it into a reliable attack — that is where the real findings come from. We also plant hidden instructions inside the content the system reads, and test whether it obeys them. Every finding is reported with the exact input, the response, and how often it reproduces.

What we look for

Whether the model can be talked out of its instructions. Whether instructions hidden in a document, ticket or web page can steer it without an attacker ever speaking to it. Whether it will reveal its own instructions, other people's data, or fragments of what it was trained on. Whether a lookup ignores the permissions of the person asking. Where it can act, whether it can be made to act for an attacker. And whether it produces harmful, unfair or confidently wrong answers.

What we will not do

AI systems do not fail predictably — the same input can work once and fail the next time, so findings are reported with how reliably they reproduce rather than as a simple yes or no. There is also no complete fix for most of these issues: a model that reads text can be influenced by text. Our advice therefore focuses on limiting what the model can reach and do, rather than on writing better instructions for it.

This tests the AI system itself. The surrounding application, its API and the infrastructure it runs on are separate services.

6 How We Rate Findings

Rated by what the attack achieves, not by how clever it is.

SeverityWhat it means
CriticalLets an attacker reach other users' data, make the system take a damaging action on their behalf, or extract credentials. Real impact outside the conversation.
HighReliably defeats your protections — exposing your instructions, reaching information the user should not see, or producing seriously harmful output.
MediumWorks, but inconsistently, or needs unusual conditions. Still worth fixing, because attacks get shared and refined.
LowMinor leakage or occasional misbehaviour with limited consequence.
ObservationBehaviour worth knowing about, or hardening advice, with no direct security impact.

Every finding records how reliably it reproduces — for example "succeeded in 8 of 20 attempts" — because that is what decides how urgently it needs fixing and lets you confirm whether your fix actually worked. Severity depends heavily on what the system can reach and do: the same jailbreak is an embarrassment on a chatbot that only answers, and a critical finding on an agent that can send email and change records.

7 Timeline

Agreed on the scoping call — not estimated here

There is no useful way to estimate this from a form. The work depends on what the system can reach and do, how many protections there are to work through, how consistently the model behaves, and how much of the effort goes into turning a partial success into a reliable attack — which is unpredictable by nature.

As a rough guide, most engagements run between one and four weeks of active testing, plus reporting. A chatbot that only answers questions sits at the shorter end; an agent that can act on your systems, reads content from outside, and holds sensitive data sits at the longer end. Once we have been through your answers above, we agree the exact duration and confirm it in writing before starting.

Save and Export Response

Your answers stay in this browser until you export or clear them.