- Permission-aware retrieval means the search itself only sees documents the user is allowed to read. The filter runs before ranking, inside the query.
- Filtering after ranking (post-filtering) leaks information through generated answers, citations and result counts, and it returns too few results when most top hits are filtered out.
- Store access control as metadata on every chunk (tenant, groups, owners) and pass the user's scope as a query filter to the vector or hybrid index.
- Test it like a security boundary: a fixture user who must never see a planted document, checked in CI.
Most retrieval-augmented generation (RAG) prototypes start with one shared index and one question: is the answer relevant? Access control gets added later, usually as a check on the results. That order is the bug.
01Why is post-filtering a RAG security risk?
Post-filtering means you search everything, take the top results, then drop the ones the user cannot see. It looks safe because the user never sees a forbidden document in the list. It leaks anyway.
The model already read it. If the filter runs after the context is assembled, or on citations only, the generated answer can contain facts from a document the user has no access to.
Scores and counts leak. “No results” versus “3 results hidden” tells a user that a matching document exists. In a sensitive domain, that alone is a disclosure.
Recall collapses. Ask for the top 10, filter out 9, and the model answers from one passage. The search did not fail. The filter threw the results away after ranking had already chosen them.
02What does pre-filtering look like?
The user's access scope becomes part of the query. The index only ranks documents inside that scope.
def retrieve(query: str, user: User, k: int = 8) -> list[Chunk]:
scope = permissions.for_user(user) # tenant, groups, explicit grants
hits = index.search(
query,
filter={
"tenant_id": user.tenant_id,
"acl": {"any_of": scope.principals}, # user id + group ids
},
top_k=k * 4, # headroom for re-ranking
)
return rerank(query, hits)[:k]Three things make this work:
- Access control lives on every chunk. When a document is split into chunks, each chunk inherits the document's tenant and access list as metadata. Chunks without access metadata are not indexed.
- The filter is applied by the index, not your code. Most vector databases and hybrid search engines support metadata filters during the search. Use them.
- Re-ranking only sees allowed chunks. A cross-encoder or model re-ranker never receives a forbidden passage, so it cannot surface one.
03Why not separate indexes per user?
Per-tenant indexes are a fine choice for hard isolation between customers. Per-user indexes are not: documents are shared across many users and groups, and copying them per user turns one permission change into thousands of writes. Use a tenant boundary for isolation, and metadata filters for everything inside the tenant.
04Permissions change. Does the index keep up?
The hard part is not the first index. It is staying correct when someone leaves a team or a folder is reshared.
- Sync access lists from the source system on change, not on a nightly batch, for anything sensitive.
- Keep the access list in a field you can update without re-embedding the text. Embeddings do not change when permissions do.
- Record the access version on each chunk, and treat a stale version as “deny” until it syncs.
Relevance is a ranking problem. Access is a filtering problem. Mixing the order of the two is how a correct-looking system leaks.
05How do you test it?
Treat retrieval as a security boundary and test it like one.
- Plant a canary document with a unique phrase, readable only by one fixture user.
- In CI, query for that phrase as every other fixture user. The expected result is zero hits and an answer that does not contain the phrase.
- Log every retrieval with the user, the scope and the chunk IDs returned, so an access review can answer “who could have seen this?”
This sits in layer L1 of how we think about systems: data and knowledge, including permissions. It is cheap to build in from the start and expensive to retrofit after a model has already answered from the wrong document.