Prompt injection
Direct instruction override, and — more dangerous — indirect injection through content the model retrieves: a PDF in the corpus, a Jira comment, a supplier email, a page fetched at runtime.
LLM assistants, retrieval pipelines over internal data, and agents wired into production tools now hold reach across your estate that no access review has ever covered. Averon tests those systems adversarially before they go live, and holds them under continuous assurance once they do — delivered in-Kingdom, inside your perimeter.
A control written into a system prompt is a request, not an enforcement point. Anything that must hold under an adversary belongs in the layer between the model and your systems.
Every document, ticket, email and web page your model reads is an instruction channel into your estate. Your ingestion pipeline is an unauthenticated input surface.
A prompt revision, a new tool registration, a re-indexed corpus, a model version bump. The system you certified in March is not the system running in May.
A conventional penetration test evaluates code paths, authentication and infrastructure. It does not evaluate what happens when a language model is given credentials, memory, and the ability to act. These are the classes Averon tests against — each with defined methodology and reportable evidence.
Direct instruction override, and — more dangerous — indirect injection through content the model retrieves: a PDF in the corpus, a Jira comment, a supplier email, a page fetched at runtime.
Regulated data, PII and credentials leaving the estate inside a completion — via retrieval over-scope, context bleed between sessions, or rendered markdown and callback URLs.
Chained tool calls that compose into an action no single call authorises. Confused deputy patterns. Unbounded loops against rate-limited internal APIs. Actions taken without a human in the loop.
Provenance of weights, adapters and embeddings. Third-party plugin and MCP server trust. Inference dependencies pulled from public registries into a production path.
The highest-severity class we report, and the most consistently present. The AI system authenticates as itself, not as the user asking. Entitlements collapse at the boundary.
A retrieval pipeline indexes a corpus using a service account that can read everything. An agent calls an internal API with a token scoped to the application, not to the person who asked. The model sits between a population of users with differentiated entitlements and a set of systems that only ever see one identity.
The result is a universal read primitive. Any user who can address the assistant can, with sufficiently patient phrasing, reach data they were never entitled to. No control is bypassed. No exploit is used. The authorization decision simply was never made.
It is invisible to your existing stack. There is no anomalous login, no malformed request, no data-loss signature. Traffic is a well-formed, authenticated API call from a trusted internal service. Your SIEM sees a healthy system. Your regulator will not.
Every finding we issue is reproducible. It carries the exact input sequence, the observed system response, the identity context under which it occurred, and the control that should have stopped it. Your engineering team can replay it. Your audit function can file it.
Two engagement modes, designed to be run in sequence. The first produces a launch decision your board can sign. The second keeps that decision true as the system changes underneath it.
A time-boxed, full-scope adversarial engagement against the system as it will actually run: real corpus, real tool registry, real identity plumbing. Concludes with a written position on whether the deployment is fit to launch.
A standing programme that re-tests on change and on schedule, so the assurance statement you hold is about the system running today — not the one assessed at launch.
Adversarial testing of an AI system requires access to the corpus, the tool registry, and production-equivalent identity. That is precisely the access a foreign vendor cannot be granted — and the reason this work has not been done.
Testing infrastructure, evidence storage and reporting remain in-Kingdom. No cross-border transfer of test data, prompts, corpus extracts, or findings. No offshore analyst access at any point in the engagement.
Findings are mapped to NCA ECC and, for financial institutions, SAMA CSF, with personal-data exposure assessed against PDPL. Reports are drafted in the register used by the people who file them.
Guardrails trained and evaluated in English fail differently in Arabic. We test in both, including dialectal and mixed-script payloads, and deliver findings in both. بالعربية والإنجليزية
Averon does not sell a model, a platform, or a guardrail product. We have nothing to protect except the accuracy of our findings — which is the only thing an assurance function is worth.
Independence charter — AveronScoping begins with a ninety-minute technical session against your architecture. No commercial commitment, and no material leaves your environment.