Authentication and application security
Multi-Tenant Boundaries: What a SaaS Architecture Should Make Impossible
A practical architecture review of authorization, tenant-scoped data, jobs, caches, storage, audit trails, and containment in multi-tenant SaaS.
A multi-tenant system is safe when an engineer can make an ordinary mistake without turning it into a cross-customer incident. Code review and testing matter, but neither should be the only thing standing between a forgotten filter and another tenant's records. The architecture should make unsafe access difficult to express, easy to detect, and limited in impact when a defect still escapes.
Products such as Cepra and Custom Name Domain make these boundaries concrete. A shared application may serve many organizations while coordinating users, domains, subscriptions, jobs, credentials, and provider resources. The tenant identifier follows work through HTTP requests, database queries, queues, caches, object storage, logs, and administrative tools. If any layer treats it as optional context, the isolation promise is weaker than the UI suggests.
The right design is not one universal tenancy pattern. Some systems use a shared schema, some separate schemas, and some separate databases. The useful question is what the chosen architecture makes impossible by construction, what it detects at runtime, and what remains dependent on developer discipline.
Establish tenant context once, then verify it repeatedly
Authentication identifies a principal. Authorization decides whether that principal may perform a specific action on a specific resource for a specific tenant. A valid session is not a substitute for that decision. OWASP recommends denying by default and validating permissions on every request because a single missed path can compromise confidentiality or integrity.
Resolve tenant context at a trusted boundary from the authenticated membership and route intent. Do not accept a tenant ID from a request body and treat it as authority. A user may choose among organizations they belong to, but the server must prove that membership, its status, and the required permission. Produce a typed request context containing principal ID, tenant ID, membership ID, role or capabilities, and a trace ID.
Then revalidate at the domain boundary. Middleware can reject unauthenticated requests and malformed tenant selection, but it cannot understand every resource relationship. A service that updates a mailbox must verify that the mailbox belongs to the selected tenant and that the membership permits the action. A job that renews a domain must prove that the job tenant, domain tenant, and provider-account tenant agree.
Avoid authorization that relies only on guessed identifiers being hard to discover. Resource IDs can leak through logs, browser history, analytics, emails, or integrations. The query must combine resource identity with tenant scope, and a missing result should fail closed.
Make unscoped database access conspicuous
In a shared-schema design, every tenant-owned row needs a non-null tenant key. Relationships between tenant-owned tables should preserve that identity. A project referencing a customer only by customer ID may allow an accidental cross-tenant association if the database has no way to assert both belong to the same tenant.
Prefer repository or data-access functions that require tenant context. Instead of a general findUser function with an optional tenant parameter, expose findTenantUser with a required tenant ID and a separately named platform lookup for rare owner operations. The safe path should be shorter and more common than the privileged path.
Composite uniqueness often needs the tenant key. A slug, provider customer reference, or external domain record may be unique within one tenant rather than globally. Constraints such as tenant ID plus slug communicate the actual invariant and prevent races that application checks cannot. Composite foreign keys can also ensure related records share a tenant where the database design supports them.
PostgreSQL row-level security provides a defense layer that can restrict which rows a database role can read or modify. When row security is enabled without an applicable policy, PostgreSQL uses default deny. That is valuable, but it requires disciplined connection context, roles, policy testing, and awareness that owners and privileged roles can bypass policies. RLS should reinforce scoped application access rather than excuse unscoped repositories.
Scope writes, transitions, and uniqueness together
Read isolation receives attention because data exposure is visible, but cross-tenant writes can be equally damaging. Every update and delete should include tenant scope in its predicate. A preliminary read followed by an update using only the resource ID reopens the boundary, especially under concurrency. Express ownership in the write itself and verify the affected row count.
Business transitions need tenant checks across every participant. Moving an invoice to paid may touch a subscription, provider account, ledger, and notification preference. Confirm those records share tenant identity before the transaction changes state. If an internal model can reference a resource from another tenant, add a constraint or validation that makes the invalid relationship impossible to persist.
Idempotency keys and deduplication keys also need tenant scope. If a client-generated key is unique globally, one tenant might accidentally or maliciously collide with another. If it is unique only by operation type without tenant identity, a retry can resolve to the wrong operation. Use a server-owned composite identity that includes the tenant and business resource.
For privileged maintenance, make the escape hatch explicit. A platform administrator may need cross-tenant search or repair, but those functions should live in a separate module, require owner authorization, record a reason, and return tenant-labeled data. Normal services should not gain a boolean bypass flag that any caller can flip.
Carry tenant identity through background jobs
Queues and scheduled work are common isolation gaps because there is no browser session to reconstruct context. Every tenant-owned job should include a stable tenant ID and the resource IDs it expects. The worker must load those resources under that tenant, not look up the resource globally and trust the payload.
Treat job payloads as untrusted historical messages. Memberships may be revoked, resources may move through lifecycle states, and permissions may change after enqueue. System jobs often run under service authority rather than user authority, but they still need current tenant and state checks. User-delegated jobs may also need the initiating principal for audit, while authorization is re-evaluated according to policy.
Partition concurrency controls by tenant and resource. A lease called renew-domains is too broad if it serializes all customers, while a lease based only on domain ID may lack tenant evidence. A key such as tenant, operation, and resource provides clearer ownership. Dead-letter queues must retain tenant labels so operators can investigate without mixing records.
Batch jobs require special care. Query one bounded tenant set at a time or return records with explicit tenant IDs and group processing. Never keep an ambient tenant variable that changes inside a concurrent loop. Pass an immutable context into each task and include it in structured logs.
Treat caches as another shared database
Caches can leak data even when the primary query is perfectly scoped. A key such as user-profile plus user ID might be safe only if user IDs are globally unique and the cached representation is not tenant-specific. Query-result caches based on a URL can collide when tenant context comes from a header or session rather than the URL.
Define a canonical cache namespace containing environment, tenant, resource type, resource identity, and a version where needed. Centralize key construction rather than assembling strings throughout the codebase. Cache tags and invalidation keys must carry the same scope; otherwise a tenant update can evict or refresh another tenant's data.
Do not cache an authorization decision longer than the membership evidence that supports it. A disabled user should not retain access because a positive permission result lives for an hour. Short lifetimes, membership-version keys, or explicit invalidation can bound that risk. Negative results also need thought because a newly granted permission may remain hidden.
Server rendering adds another cache boundary. If a framework can statically or globally cache a response that depends on session or tenant headers, make the route dynamic or include the correct variation key. Review full-route, data, client-router, and CDN caching separately; they do not necessarily share invalidation behavior.
Namespace files, secrets, and provider resources
Object storage should make tenant ownership visible in both metadata and path. A prefix containing tenant ID helps lifecycle policies and incident analysis, but path structure alone is not authorization. Generate signed access only after checking the database ownership record. Never let a client request an arbitrary storage key and receive a signed URL.
Store the expected content type, size, hash, uploader, tenant, and business resource. On download or processing, compare storage metadata with the database record. A background scanner should receive tenant context and write results back through a scoped predicate. Deletion jobs should verify ownership again rather than trusting a stale key.
Provider resources also need mappings. A payment customer, registrar account, DNS zone, or mailbox organization must be associated with one tenant locally. Before using a provider object ID supplied by the browser or webhook, confirm it matches the tenant's mapping. Webhook signatures prove the sender, not that the referenced object belongs to the local tenant being updated.
Credentials should be isolated by tenant when the provider model requires it, encrypted at rest, and never copied into job payloads or logs. Workers load credentials through an authorized mapping at execution time. Rotation then changes one controlled record rather than invalidating a queue of messages containing old secrets.
Build audit trails that preserve subject and scope
An audit event needs more than actor and action. Include tenant, subject resource, result, request or job identity, authorization basis, important before-and-after state, and a timestamp. For support actions, record the operator's platform identity and the tenant they intentionally entered. Avoid logging secrets or unnecessarily duplicating personal data.
Security logs should make cross-tenant anomalies detectable. Examples include a resource ID found under a different tenant, repeated denied lookups, mismatched provider mappings, a job payload whose tenant disagrees with its resource, or a privileged query without a support case reference. These are signals of both attacks and implementation defects.
Keep audit storage append-oriented and restrict who can read it. Tenants may receive a filtered activity history, while platform security retains a broader view. The filtering code itself must be tenant scoped. Export jobs and analytics pipelines are not exempt from isolation simply because they run offline.
Tracing should propagate tenant identity in controlled attributes, not in public error text. Cardinality and privacy policies may prohibit using raw tenant IDs in every metric label; logs and traces can hold a stable internal identifier while aggregate metrics measure denied access and mismatches without exploding label counts.
Contain failures with deliberate deployment topology
Shared infrastructure creates a larger blast radius, so quotas and fairness are part of tenancy. One tenant's expensive report, webhook storm, or failed integration should not exhaust workers for everyone. Apply per-tenant rate limits, queue concurrency, storage quotas, and provider budgets. Reserve capacity for control-plane and recovery operations.
Database-per-tenant or schema-per-tenant designs can strengthen some boundaries but add migration, connection, and operational complexity. Shared-schema designs can be appropriate when repository scoping, constraints, RLS, testing, and monitoring are strong. Choose based on risk, scale, compliance, and operating capability rather than treating one topology as automatically secure.
Feature rollout can also contain failures. Enable risky migrations or integration changes for a small tenant cohort, observe invariants, and expand gradually. Keep configuration tenant scoped and audited. A global fallback should not silently expose a feature whose authorization model is incomplete.
Incident tooling must support selective isolation: pause one tenant's jobs, revoke one provider mapping, invalidate one namespace, or disable one feature without shutting down the platform. Those controls reduce pressure to make unsafe manual database edits during an outage.
Prove isolation with adversarial tests
Unit tests should assert allow and deny decisions, but integration tests must try the same identifier under two tenants. Create mirrored fixtures, authenticate as tenant A, and attempt to read, update, delete, export, cache, enqueue, download, and subscribe to tenant B's resources. Test positive and negative paths for every transport, including server actions, route handlers, sockets, and background workers.
Add static conventions where possible: tenant-owned repository methods require a context type, tenant columns are non-null, storage access goes through one service, and privileged clients have visibly different names. Review migrations for new tables missing tenant identity and indexes. Test RLS with the same database roles used in production, not a superuser that bypasses policy.
Failure tests matter too. Revoke membership after a job is queued. Reuse an idempotency key across tenants. Poison a cache entry under the wrong namespace. Deliver a valid provider webhook for a resource mapped to another tenant. Run concurrent updates against resources with similar IDs. The expected result is denial, quarantine, or a contained error with audit evidence.
The architectural standard is stronger than saying developers must remember tenant filters. A mature SaaS makes the scoped path natural, the unscoped path privileged, invalid relationships unpersistable, asynchronous context explicit, and suspicious mismatches observable. That is how tenant isolation becomes a system property rather than a coding convention.
Primary sources
- 1.Authorization Cheat Sheet — OWASP
- 2.PostgreSQL Row Security Policies — PostgreSQL
- 3.PostgreSQL CREATE POLICY — PostgreSQL
- 4.Next.js Data Security Guide — Next.js
Portfolio evidence
CEPRA
Established organization-safe operations and connected delivery, billing, capacity, and executive reporting across the web platform.
View case studyCustom Name Domain
Delivered the central commerce and branded-email workflows needed to move multi-vendor setup into a repeatable self-service product path.
View case studyRelated writing