How to Prove One Tenant Can't Read Another Tenant's Data in Supabase

Your tenant isolation lives in hundreds of migrations. Nobody can see all of it, and nobody can say for sure that it works. Here is how to get that certainty back without rewriting anything.

The question nobody can answer with confidence
Every multi-tenant Supabase project starts the same way.
There is one organizations table, a members table, and a handful of Row Level Security policies that compare org_id with something from auth.uid() or auth.jwt(). It is small enough to keep in your head. It works.
Then the product grows.
A SECURITY DEFINER function is added to make an invite flow work. A view is created to simplify a dashboard query. A Storage bucket gets its own policies on storage.objects. An Edge Function needs the secret key, which bypasses RLS, to process a Stripe webhook. A Next.js API route takes a projectId from the URL. Someone adds a permissive policy while debugging and forgets about it.
Each change is reasonable on its own. Each one lands in its own file in supabase/migrations.
A year later, the layer that decides who can see which data is spread across hundreds of migrations. Nobody on the team can see all of it at once. And when a customer, an investor or an auditor asks the simplest question, "Can one tenant ever read another tenant's data?", the honest answer is: "We think not."
That is the pain. Not a known vulnerability. A lack of visibility, and no way to be sure the data is actually protected.
It is also the question buyers now ask before they sign. In one recent audit request, a team extending their Supabase platform from one company to several listed its top priority as "confirming one company's data can never be read from another's session." Another founder, about to open their app to twenty external users, asked for "one automated test that signs in as user A and proves user B's data is unreachable, so I can run it before every deploy."
And the worry is justified. Supabase's own documentation lists the doors that stay open even when every table has RLS enabled:
Views. "Views bypass RLS by default because they are usually created with the postgres user." Imagine invoices is protected so each organization sees only its own rows, and someone adds create view invoice_summary as select * from invoices to speed up a dashboard. A member of another organization who queries invoice_summary gets every organization's invoices, because the view reads the table as its creator. The fix is one line, security_invoker = true, but only if someone notices.
User metadata. raw_user_meta_data "can be updated by the authenticated user", so "it is not a good place to store authorization data." If a policy reads the tenant from user_metadata, a user can change their own tenant with supabase.auth.updateUser() and walk into someone else's data.
Security definer functions. They run with the privileges of their creator, and "a security definer function in an exposed schema is callable over the Data API with the creator's privileges."
The secret key. It "authorizes access through the service_role Postgres role, which has the bypassrls attribute." Every route or Edge Function that uses it must do its own tenant check.
None of these look wrong when you read a single migration. They only show up when someone tries to cross the line.
Not sure your tenant data is really isolated today? Our Supabase Security & Tenant Isolation Audit checks your RLS policies, views, functions, Storage and server routes, and ends with a Cross-Tenant Isolation Verdict.
Not ready yet? Start with a free security scan of your app's public surface.
The low-hanging fruit is not a rewrite

When teams feel this pain, the instinct is to go big.
Review every RLS policy, function and grant. Redesign the schema. Move to schema-per-tenant or project-per-tenant. Introduce a new, "correct" way of organizing policies that gives the best security and the best performance.
All of that may be worth doing one day. It is not where you start. It takes weeks, it touches everything, and at the end you still have the same question: does it actually work? We saw this in a Supabase RLS performance case study: policies form one combined authorization model, so "changing one condition can affect both performance and security semantics." A rewrite without a safety net is a new risk, not a fix.
The pragmatic first step is to test the behaviour, from the outside, the way an attacker would. Sign in as a member of Tenant B, try to reach Tenant A's data through every door your app has, and check that every door is closed.
That is what end-to-end (E2E) tests do, and Playwright is the tool we use for them. Here is why it fits this pain so well.
It tests the whole path, not just the database. Database tests with pgTAP are valuable, and Supabase's own testing guide shows a multi-tenant example built with them. But pgTAP only sees Postgres. Many tenant leaks happen outside it: an API route that uses the secret key and filters only by an ID taken from the URL, an Edge Function, a Storage download. Supabase's testing guide puts it simply: "Testing through application code provides end-to-end verification." A Playwright test goes through the same Next.js route, the same supabase-js client, the same PostgREST API and the same Storage endpoint your users go through.
It doesn't care how isolation is implemented. Playwright's best practices tell you to "test user-visible behavior", not implementation details. For tenant isolation, that is exactly the point. Whether access is enforced by an RLS policy, a JWT claim in app_metadata, a membership table or a server-side check, the test asks only one thing: did Tenant B see Tenant A's data? You can refactor every policy underneath, and the test still tells you the truth.
It follows the official advice on what to test. Supabase's testing guide says to "test with different user roles: anonymous and authenticated" and to "always test negative cases: what users should not be able to do." Its RLS guide goes further for shared tables: assert that a member who is not the owner can do what the policies allow, "and that a non-member cannot." Security testing frameworks built on Cucumber have used the same three checks for years: an authorised user can reach their own resource, a different user cannot (CWE-639, authorization bypass through a user-controlled key), and an unauthenticated visitor cannot (CWE-306).
Roles are built in. Playwright's authentication guide shows how to sign in once per role and save each session with storageState, and even run several signed-in roles in one test, each in its own browser context. "Member of Tenant A" and "Member of Tenant B" side by side is a standard pattern, not a hack.
It runs on every pull request. The risk in a growing Supabase project is not the policy you wrote last year. It is the migration someone merges next week. Supabase recommends running tests "automatically on every pull request", and Playwright has a guide for running tests on CI. A cross-tenant suite in CI turns "we think not" into a green or red check on every change.
It answers the question in the language people ask it. A failing test named "Member of Tenant B cannot open Tenant A's invoice" needs no explanation for a founder, a customer or an auditor.
It is not magic. E2E tests are slower than database tests. They cannot roll back with a transaction, so Supabase advises designing them with "unique user IDs for each test case" instead of relying on a clean database. And most importantly, an E2E suite only proves the scenarios you wrote. If nobody wrote down that an Editor must not change billing, nobody will test it.
That is why the next step happens before any test is written.
Good coverage starts before the test: the artifacts
You can ask an AI to "write Playwright tests for tenant isolation". You will get tests. You will not know what they cover.
Playwright itself points in a better direction. Its Test Agents split the work into three roles: a planner that "explores the app and produces a Markdown test plan", a generator that "transforms the Markdown plan into the Playwright Test files", and a healer that repairs failing tests. The planner takes context: "a seed test that sets up the environment necessary to interact with your app" and, optionally, "a Product Requirement Document (PRD)."
In other words, the quality of the tests depends on the quality of what you give the AI before it starts. For tenant isolation, these are the artifacts that matter, in order of importance.
1. User stories with clearly defined actors. This is the most important one. In our fractional CTO case study, the first step was to identify "the actors, roles, concepts, events, rules and relationships" of the business, before a single feature was built. Tenant isolation is entirely about who. Who is the owner of Tenant A? Who is a regular member? Who is a member of Tenant B? Who is an anonymous visitor? Who is a platform admin, and what may they see? If the actors are vague, the tests will be vague.
2. Functional requirements. The rules each actor lives by. "An Editor can publish posts but cannot manage members." "A Viewer can see invoices but not prices." These become the expected results.
3. An information model. Which tables, views, functions and Storage buckets hold tenant data, and how each one knows which tenant a row belongs to. The more of this the AI has, the fewer doors it misses.
These artifacts must be kept in a format with a strong structure. Strong means easy to maintain and easy to follow, for a human and for an AI. That format is Gherkin.
Cucumber, the project behind Gherkin, describes it as a set of keywords that "give structure and meaning to executable specifications." Each part has a job:
Feature groups the scenarios for one capability, such as Tenant data isolation.
Rule captures one business rule, such as Members only see their own organization's data.
Background sets up shared context, such as the two organizations and their members.
Scenario Outline with Examples runs the same scenario for many actors and resources. For permissions, this is the key feature: one table becomes your access matrix.
Given / When / Then puts the system in a known state, describes the action, and states an outcome that is observable from outside the system.
Cucumber's own guidance is to "describe the intended behaviour of the system, not the implementation". Write "a member of Tenant B opens Tenant A's project", not "click the third row in the table". Declarative scenarios survive UI changes, and they read as living documentation.
If your team needs a way to find these rules, Cucumber's Example Mapping helps: a story card, rule cards beneath it, example cards under each rule, and question cards for what nobody can answer yet. Each rule card becomes a Gherkin Rule. Each example becomes a row in your Examples table.
We described the general workflow, Gherkin as the source of truth and Playwright tests as generated artifacts, in From Gherkin to Playwright: AI-Generated Tests for Supabase Applications. The team owns the specification; AI owns the generated artifact. Tenant isolation is where that idea pays off most.
What it looks like in practice

How others write permissions in Gherkin
We are not the first to describe access control this way.
BDD-Security, an open-source security testing framework built on Cucumber, ships an authorisation.feature with three Scenario Outlines: users "can view restricted resources for which they are authorised", users "must not be able to view resources for which they are not authorised" (tagged CWE-639), and "un-authenticated users should not be able to view restricted resources" (CWE-306). Its trick is simple: record what Bob can see, then replay the same requests as Alice and check that Bob's data is not in any response.
Playwright gives you the two halves of a role matrix: one saved session per role, and API requests sent from inside a UI test to check what the server really returns, not only what the page shows.
Supabase's own pgTAP guide tests a multi-tenant publishing app with owners, admins, editors and viewers, with "cross-organization data isolation" as a focus area, and expects forbidden actions to fail with SQLSTATE 42501.
Here are three examples that put these patterns into the words of a typical Supabase app.
Example 1: Members never see another organization's rows
Feature: Tenant data isolation
Background:
Given organization "Acme" with owner "Alice" and member "Adam"
And organization "Globex" with owner "Gina"
And "Acme" has a project "Acme Roadmap" and an invoice "ACME-001"
Rule: Members only see their own organization's data
Scenario Outline: <actor> opens a resource of Acme
Given "<actor>" is signed in
When they open the <resource> "<name>"
Then they <outcome>
Examples:
| actor | resource | name | outcome |
| Alice | project | Acme Roadmap | see it |
| Adam | project | Acme Roadmap | see it |
| Gina | project | Acme Roadmap | get "not found" |
| Gina | invoice | ACME-001 | get "not found" |
Scenario Outline: An anonymous visitor cannot read tenant data
Given nobody is signed in
When they request the <resource> "<name>"
Then no data is returned
Examples:
| resource | name |
| project | Acme Roadmap |
| invoice | ACME-001 |One table, three kinds of actor: the owner, a member, a member of another tenant, plus the anonymous visitor. Adding a new table to the information model means adding a row, not writing a new test.
Example 2: Writes cannot cross the line
Reads are only half of isolation. Supabase's RLS guide asks for a separate policy for each operation, select, insert, update and delete, and each one is a separate door to test.
Rule: Members can only write inside their own organization
Scenario: A member cannot join another organization on their own
Given "Gina" is signed in
When she tries to add herself as a member of "Acme"
Then the request is rejected
And "Acme" still has 2 members
Scenario: A member cannot move a project to another organization
Given "Adam" is signed in
When he tries to change the organization of "Acme Roadmap" to "Globex"
Then the request is rejected
And "Acme Roadmap" still belongs to "Acme"Example 3: Files in Storage follow the same rules
Supabase Storage is secured with RLS policies on storage.objects, and "by default Storage does not allow any uploads to buckets without RLS policies." That makes Storage easy to forget in a test plan, and Storage is where invoices, contracts and exports live.
Rule: Files belong to the organization that uploaded them
Scenario Outline: <actor> downloads Acme's contract
Given "Alice" uploaded "contract.pdf" to the "documents" bucket for "Acme"
And "<actor>" is signed in
When they download "contract.pdf"
Then they <outcome>
Examples:
| actor | outcome |
| Adam | receive the file |
| Gina | are refused |What an AI-generated Playwright test can look like
Given Example 1, a generator that follows Playwright's authentication guide produces something like this. The setup project signs in each actor once and saves their session. The test then checks both doors, the page and the API, because a hidden page does not prove that the data is protected.
// tests/tenant-isolation.spec.ts (generated from features/tenant-isolation.feature)
import { test, expect } from '@playwright/test';
import { createClient } from '@supabase/supabase-js';
import { seed } from './fixtures/seed'; // ids created by the seed test
const supabaseFor = async (email: string) => {
const client = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_PUBLISHABLE_KEY!);
await client.auth.signInWithPassword({ email, password: process.env.TEST_PASSWORD! });
return client;
};
test.describe('Tenant data isolation', () => {
test.describe('Gina (Globex) opens resources of Acme', () => {
test.use({ storageState: 'playwright/.auth/gina.json' });
test('project "Acme Roadmap" is not found in the app', async ({ page }) => {
await page.goto(`/projects/${seed.acmeRoadmapId}`);
await expect(page.getByRole('heading', { name: 'Not found' })).toBeVisible();
await expect(page.getByText('Acme Roadmap')).toHaveCount(0);
});
test('project "Acme Roadmap" is not returned by the API', async () => {
const gina = await supabaseFor('gina@globex.test');
const { data } = await gina.from('projects').select('*').eq('id', seed.acmeRoadmapId);
expect(data).toEqual([]); // RLS filters the row out; it does not raise an error
});
test('invoice "ACME-001" is not returned by the API', async () => {
const gina = await supabaseFor('gina@globex.test');
const { data } = await gina.from('invoices').select('*').eq('number', 'ACME-001');
expect(data).toEqual([]);
});
});
});Each Scenario Outline row becomes a test, and each test name reads like the Gherkin it came from. When one turns red, everyone knows which promise was broken.
Benefits and downsides
Benefits
Visibility. The feature files are a readable map of who can access what. That is the map that was missing from the migrations.
Coverage you can count. Every actor × resource row is a check. Gaps are visible as missing rows, not hidden in code.
Cheap to maintain. When the UI changes, you regenerate the tests. The specification stays the same.
Reusable beyond security. The same Playwright journeys can later drive browser-based load tests with Artillery and Supabase tracing, so the investment pays off twice.
Regression safety. In CI, a migration that opens a door fails the build before it reaches production.
Downsides
Generated tests are not automatically correct. Playwright's docs say generated tests "may include initial errors", and the healer may return "a skipped test if the healer believes that functionality is broken." For a security test, a silently skipped test is worse than none. Every generated test needs a human review, and skipped tests must fail the build.
AI tends to test what it sees. A test that checks a hidden button proves nothing about the data. Require an API-level assertion for every negative case; Playwright supports validating server-side postconditions from a UI test.
It proves only what is written. The suite is exactly as good as the user stories, actors and information model behind it.
It needs a controlled environment. Seeded staging data, dedicated test users, and saved sessions in playwright/.auth that must never be committed. Playwright warns that they "may contain sensitive cookies and headers that could be used to impersonate you."
Closing the loop
There is one more benefit, and it goes beyond testing.
In a full software development lifecycle, the requirements written at the start are the same ones used at the end, to check that the implementation does what was agreed. That is why we treat sign-off as a business decision, made before delivery begins. Gherkin makes that literal. The user stories you wrote with your stakeholders become the feature files, the feature files become the tests, and the test report becomes the answer to the question from the beginning of this post.
Not "we think not". "Here is the proof, and it runs on every deploy."
If you are not sure where your Supabase project stands today, and you want an independent answer before your next customer asks, start with our Supabase Security & Tenant Isolation Audit. It ends with a Cross-Tenant Isolation Verdict and a prioritized remediation roadmap.
Or
A passive scan of your live app plus our Security Guidebook. It checks your public pages; it does not test tenant isolation.



Comments