
RAG security is being solved at the wrong layer. The missing concept is chunk provenance — and it changes what "secure RAG" means
You've locked down your vector database. You've added input sanitization. You're filtering outputs and logging every request. Your secure RAG architecture checklist looks solid.
And yet, if someone asked you a simple question — can you prove that the chunks your model consumed in yesterday's customer-facing answer were authentic, came from an approved source, and hadn't been modified since ingestion? — you couldn't answer it. Not with evidence. Not with anything an auditor or an incident-response team would accept.
That's not a gap in your controls. It's a gap in the model of the problem. Enterprise RAG security today is built entirely around inspection —inspecting inputs, inspecting outputs, inspecting traffic. What it's missing is provenance: a verifiable record of where each piece of retrieved context came from, who vouched for it, and whether it's still the same content that was vouched for.
This post makes the case for why chunk provenance is the missing primitive in RAG security, and what a provenance-first architecture looks like.
Ask most AI architects how they've secured their RAG application, and you'll get a predictable list: sanitize user inputs before retrieval, filter model outputs before delivery, lock down access to the vector database, log requests and responses for audit.
All of these are necessary. All of them share the same blind spot: they treat retrieved content as trustworthy by default. Once a chunk is in the knowledge base, every downstream control assumes it belongs there and says what it said when it was ingested.
Neither assumption is safe. And no amount of input or output filtering can validate them, because filters inspect what content looks like, not where it came from or whether it changed.
The right mental model for enterprise RAG isn't an application with a database behind it. It's a supply chain: content flows from source systems (SharePoint, Confluence, S3, SQL, APIs) through an ingestion pipeline (chunk, embed, index) into a vector store, gets retrieved at query time, and lands in the model's context.
Software engineering already learned this lesson the hard way. After a decade of supply chain attacks, the industry converged on artifact signing, transparency logs, and frameworks like SLSA: you don't trust a dependency because it passed a scan, you trust it because it carries a verifiable signature tracing it to an authenticated builder.
The RAG supply chain has none of this. Every handoff — source to pipeline, pipeline to vector store, vector store to model context — is unauthenticated. There is no cryptographic anchor between any two stages. That structural weakness shows up as four distinct failure modes:
An attacker doesn't need access to your application; they need to get a malicious document into any ingestion path. Research presented at USENIX Security 2025 (PoisonedRAG) showed that injecting as few as five doctored documents into a corpus of millions achieves roughly 90% attack success. The OWASP Top 10 for LLM Applications ranks data poisoning among the highest-risk categories for a reason: the attack is persistent, hard to detect, and amplified every time the poisoned chunk is retrieved.
The sharper edge of poisoning. Malicious instructions embedded in a retrieved chunk arrive in the model's context alongside legitimate content, and the model has no way to distinguish them. Sanitization helps at the margins, but adversarial content evolves faster than filter rules — and a filter can't tell you whether the chunk had any business being in the corpus at all.
Anyone — or any compromised credential —with write access to the vector store can modify chunk content after ingestion . No integrity check exists at retrieval time, so the model consumes the altered content and nobody notices. This is the insider-threat and compromised-pipeline scenario that perimeter controls structurally cannot see, because the tampered chunk flows over the same internal paths as an authentic one.
The non-adversarial failure mode, and often the most common one. The source document gets updated; the vector store copy doesn't. The model confidently answers from last quarter's policy, yesterday's price list, the superseded clinical guideline. Nothing was attacked. The answer is still wrong, and nothing in the stack flags it.
Notice what these four have in common: none of them is detectable by inspecting the chunk's content. A poisoned chunk reads like a normal document. A tampered chunk reads like a normal document. A stale chunk is a normal document —just the wrong version. The only way to catch any of them is to know the chunk's history: who ingested it, when, from what source, and whether the bytes have changed since.
That history has a name. It's provenance.
Chunk provenance means every unit of retrievable context carries a cryptographic record of its origin and integrity— created at ingestion, verified at retrieval, and preserved as evidence afterward. Concretely, it decomposes into three operations at three points intime:
When a chunk enters the corpus, compute a digest of its content and sign it — not with a long-lived key sitting in pipeline config, but with a short-lived certificate bound to an authenticated identity: the ingestion pipeline's workload identity for routine content, or a human curator's SSO identity for high-stakes corpora. Record the signing event in an append-only transparency log, and store the signature bundle alongside the chunk in the vector store metadata. This is the Sigstore model from software supply chain security, applied to retrieval units instead of build artifacts. Signing once per chunk, asynchronously, costs nothing on the request path.
At request time, before retrieved chunks enter the model's context, recompute each chunk's digest and compare it against the signed digest in its stored bundle. A local hash match takes under a millisecond per chunk and requires no network call — the transparency log stays off the hot path, held in reserve as the canonical proof if a signature is ever challenged. Any tampering after signing produces a mismatch, deterministically. Layer a trust policy on top: which signers are acceptable, how fresh a signature must be, and what happens on failure (fail closed for regulated query classes, log-and-allow for research ones). Not every corpus deserves equal trust — a curator-approved corpus of regulatory guidance and a feed of scraped public content should not be interchangeable in the model's context, and trust tiers make that difference enforceable rather than aspirational.
For every inference request, emit a signed manifest — an Assembly Manifest — recording exactly which chunks entered the context, each one's signer, trust tier, and verification result, the policy decision, and hashes binding the query and the model's response to that verified context. This artifact is the audit-grade answer to "what did the model see?" Application logs can't provide it; logs are reconstructions, and reconstructions aren't evidence. A signed manifest anchored to a transparency log is.
Sign, verify, audit. Ingestion time, request time, and continuously. That's the whole model — and none of it requires replacing your vector store, your ingestion pipeline, or your model endpoint. Provenance is a layer that wraps the supply chain you already run.
If you're building or auditing a RAG architecture, here's the layer-by-layer model that holds up once provenance is treated as the backbone rather than an afterthought:
Layer1 — Source and ingestion controls.
Approved connectors, source classification, and content hashing at ingest. This is where provenance is created; everything downstream depends on it.
Layer2 — Chunk signing.
Identity-bound signatures per retrieval unit, recorded in a transparency log. The step almost every RAG deployment today skips entirely.
Layer3 — Retrieval-time verification and policy.
Hash-match verification plus trust-tier policy enforcement before context reaches the model. This is the control point that catches poisoning, tampering, and drift in the milliseconds between retrieval and generation.
Layer4 — Per-request evidence.
Signed Assembly Manifests written to tamper-evident storage, with retention aligned to your regulatory regime.
Layer5 — Inspection controls.
Prompt firewalls, output filtering, access control, rate limiting. Still valuable — but as the final layer, complementing provenance rather than substituting for it.
Most enterprise RAG security investment today goes into Layers 1 and 5, because those map to disciplines security teams already know: ingestion hygiene and perimeter inspection. Layers 2 through 4 —the provenance core — are where the four failure modes actually get closed, and where almost nobody is standing.
Regulators have converged on the same question this post opened with. EU AI Act Article 12 requires tamper-evident, automatically generated records for high-risk AI systems, with full enforcement in August 2026 and penalties reaching EUR 35M or 7% of global turnover; Article15 explicitly requires resilience against data poisoning. NIST AI RMF calls for documented data lineage and integrity controls. Sector regimes — the US Treasury's AI framework for financial services, 21 CFR Part 11 in life sciences, Colorado SB 205 — stack on top, and all of them reduce to the same operational demand: prove what your AI system saw, and prove nobody changed it.
Reconstructed application logs don't meet that bar. Cryptographic provenance does — and the compliance budget already exists to pay for it, because provenance evidence is precisely the artifact those audits consume.
This is the gap our RAG Provenance Platform is built to close. It implements the full provenance model described above — identity-bound signing at ingestion via Sigstore-compatible infrastructure (Fulcio, Rekor, Red Hat Trusted Artifact Signer), sub-milli second hash-match verification at retrieval, per-corpus trust tiers and policy enforcement, and a signed Assembly Manifest plus query able provenance ledger for every request. It deploys inside your AWS environment, wraps the vector store and ingestion pipeline you already run, and produces the evidence your auditors actually accept — with verification overhead validated at under 50msend-to-end across real Bedrock and frontier API endpoints.
If you're building or reviewing enterprise RAG and want to see what verifiable retrieval looks like against your own corpus and your own policies, talk to the Traction Layer team about our Design Partner Program.
.jpg)
The whole shared-infrastructure, multi-tenant thing is basically the hero of the startup world. It's that super lean, budget-friendly engine that gets tons of products from an idea to their first users.
But if you're an ambitious B2B SaaS company working with AI, there's a catch. The very setup that helped you get started starts to become a major liability, like a boat anchor slowing you down when you try to land those big-money enterprise clients.

Sure, sharing infrastructure is fine when you're just messing around, but if you keep trying to grow on a shared system, you're not just piling up tech problems; you're actively shrinking the pool of customers who can buy from you.
It's time to shake things up. To land those six- and seven-figure deals, you have to move on from shared resources to dedicated ones. Switching to a single-tenant setup isn't a luxury; it's something you have to do.
The cracks in a shared system don't really show when you only have a handful of people testing it out. They start to appear when you're handling real work and getting looked at under the microscope of an enterprise security review. Here are three things every CTO and founder needs to face.
When you're sharing, your app shares GPU power with others. So when another user kicks off some huge, messy job, your slick, real-time AI coding assistant suddenly slows to a crawl. These random slowdowns (what a bad Time-to-First-Token looks like) are just awful for the user experience. This fight for resources makes it almost impossible to figure out if you're even making money on a per-user basis.
Enterprises in areas like finance and healthcare don't just like to have their data isolated; they demand it. A shared database and app, no matter how well you build it, always has a risk of data getting mixed up, and that's a deal-breaker. When a potential customer's request form asks for a private cloud deployment, customer-managed encryption keys, or specific data location rules to meet standards like SOC 2 or HIPAA, a standard multi-tenant platform means you're out before you even start. You're failing the security check from the get-go.
The rigid nature of a one-size-fits-all shared system means you can only innovate as fast as your slowest customer. You've built a new, better AI model, but you can't release it because one big client isn't ready for the change. Being stuck like this kills your ability to move fast, slows down new features, and makes your engineers waste time supporting clunky old code instead of creating cool new stuff.
Switching to a single-tenant approach is more than just fixing problems; it's about giving yourself a real edge over the competition. It's the foundation that lets you confidently say "yes" to what enterprise customers ask for.
A true single-tenant AI setup means giving each customer their very own, dedicated version of your app, ideally inside their own cloud space (VPC). This is the gold standard. It offers solid proof that their data is separate, lets them manage their own encryption keys, and makes it easy to have zero-data-retention policies. It's not just a feature; it's you showing them you take their data security as seriously as they do.
With dedicated GPUs for each customer, the "noisy neighbor" problem is gone. Performance becomes something you can predict and control. You can guarantee specific speeds and response times, figure out your cost-per-token down to the penny, and build a business with way better economics.
Getting from a cool product to a business that can really scale comes down to smart infrastructure choices. While a shared system gets you in the door, it won't help you win the game.
A multi-tenant setup tells them you're not ready for that kind of commitment. A secure LLM deployed in a single-tenant setup is the price of admission to the enterprise big leagues. It shows you're mature and you get what it takes to work with them.
The biggest challenge for AI startups is closing the gap between a cool prototype and a real-deal application that big companies will trust. This means you need a new philosophy for your infrastructure, one that's all about isolation, control, and getting the most performance for your money.
Traction Layer's Single-Tenant Secure Enclave architecture is built for this exact move. We give you a production-ready, high-performance inference platform that runs inside your customer's own cloud, turning security and compliance from a sales hurdle into your biggest selling point.
Check out Traction Layer AI for your VPC.
Tell us what you're running (your targets for speed, simultaneous users, context lengths, and models).
We’ll map out a plan for you to get predictable performance and the kind of unit economics you need to land your next huge enterprise deal.