Scope GraphQL penetration testing across the schema, resolvers, field-level authorization, query-cost controls, batching, mutations and downstream services. Testing only the HTTP endpoint misses the architecture that makes GraphQL powerful: one client document can traverse many objects, invoke multiple resolvers and combine data protected by different services and policies.
Define the architecture question before testing
Map the GraphQL engine, schema sources, federation gateway, resolver owners, data loaders, caches and downstream APIs. For each sensitive type and field, identify whether authorization happens at the gateway, resolver, domain service or data layer and whether parent authorization is incorrectly assumed to protect child fields.
Write the expected security invariant and the prohibited outcome in language that engineering, security and operations interpret the same way. Name the identity, trust transition, protected asset and authoritative state. That statement becomes the basis for scope, evidence and retesting.
Build the minimum evidence pack
- A current component and trust-boundary diagram for the environment under test.
- Identity, credential, network and data-flow details relevant to the decision.
- Representative accounts, workloads and synthetic records with known ownership.
- Logs or traces that follow the action through every material control point.
- Explicit safety limits, stop conditions, owners and restoration steps.
Provide schemas for enabled environments, representative roles and tenants, persisted-query rules, complexity configuration, federation topology and logs that correlate an operation with resolver and downstream calls. Include disabled introspection assumptions and any alternative schema-discovery path available to legitimate clients.
Map schema exposure and discovery
The assessment needs an authoritative view of reachable operations even when public introspection is disabled. Mobile clients, documentation, errors and persisted queries may reveal a schema that has drifted from the approved contract.
- Compare authenticated and anonymous schema visibility.
- Identify deprecated fields and mutations still reachable.
- Check federation and administrative subgraphs for accidental exposure.
- Test suggestions and validation errors for excessive disclosure.
- Verify persisted-query controls cannot be bypassed with arbitrary documents.
Report reachable operations by role and environment. Treat introspection as an exposure decision, not a substitute for field authorization.
Test resolver and field-level authorization
A parent object may be visible while one field contains privileged or cross-tenant data. Each resolver must enforce the policy required by the field and object relationship.
- Request sensitive child fields from an authorized parent.
- Use aliases and fragments to repeat or disguise protected selections.
- Traverse relationships from an owned object to a foreign object.
- Compare node or global-ID lookups with normal collection routes.
- Check nullable errors do not return partial protected data.
Capture the exact query, variables, actor, tenant and returned field path. Confirm downstream data state when a mutation is involved.
Challenge mutations and business state
Mutation authorization must cover target ownership, allowed state transition and side effects—not merely the mutation name. Nested input can hide foreign identifiers or privileged options.
- Swap nested object, owner or tenant identifiers.
- Attempt invalid state transitions and repeated submission.
- Combine allowed and prohibited changes in one mutation.
- Use aliases to invoke the same side effect repeatedly.
- Verify upload, import and webhook-triggering mutations enforce equivalent policy.
Inspect authoritative state and side effects, because GraphQL errors may coexist with partial execution. Record which resolver committed each change.
Assess depth, breadth, aliases and batching
Query complexity controls should represent actual backend work. Simple depth limits miss wide fragments, aliases, expensive filters, nested pagination and batched operations.
- Increase breadth while keeping depth constant.
- Repeat costly fields through aliases and fragments.
- Combine many operations in one HTTP request where supported.
- Vary pagination bounds, filters and nested connections.
- Test whether rejected operations still trigger partial resolver work.
Use conservative agreed limits and monitor resolver counts, latency and downstream calls. This is resilience validation, not uncontrolled load testing.
Verify data loaders, caches and tenant keys
Batch loaders and caches can collapse requests across users or tenants. A resolver may authorize correctly yet return a value cached under an incomplete key.
- Alternate equivalent queries between two tenants.
- Request the same global ID under different roles.
- Test per-request and shared loader boundaries.
- Check authorization changes invalidate cached fields.
- Vary locale, permission and preview modes that affect representation.
Prove the source of the returned value using canary records and trace identifiers. Separate stale legitimate data from true cross-context reuse.
Test federation and downstream trust
A federation gateway composes fields from subgraphs that may trust forwarded identity differently. Direct subgraph routes or entity-resolution calls can bypass gateway policy.
- Attempt agreed direct access to subgraphs.
- Remove or forge gateway identity headers.
- Test entity references containing foreign tenant identifiers.
- Compare authorization across duplicate fields owned by different subgraphs.
- Verify service tokens cannot turn user requests into unrestricted backend access.
Trace the operation from gateway plan to subgraph and authoritative store. Identify the layer that must own the durable policy.
Execute as controlled hypotheses
- Establish the legitimate baseline and capture authoritative state.
- Change one trust variable—identity, route, scope, object, time or environment.
- Observe the decision at each layer rather than relying only on the response.
- Stop at the minimum proof that demonstrates or rejects the hypothesis.
- Restore test state and record residual uncertainty or blocked coverage.
Use operation allowlists, synthetic objects and explicit cost ceilings. Coordinate any batching or complexity tests with operations and stop before availability impact. Exercise both legitimate client patterns and minimally altered negative cases so results reflect the deployed architecture rather than an artificial scanner workload.
Report the architectural consequence
Organize findings by failed invariant—schema exposure, field authorization, state transition, cost control, cache context or subgraph trust. Include query, variables, selected path, actor, tenant, resolver chain and final effect. Avoid calling every issue an endpoint flaw when the cause is shared resolver or federation policy.
Separate observed evidence from inference. State prerequisites, repeatability, blast radius and the shared component responsible for the decision. When multiple findings have one architectural cause, keep the individual proofs but group the remediation around the common control.
Required closure evidence
- The original proof no longer succeeds.
- A legitimate workflow still succeeds under the intended identity and route.
- Alternate consumers of the same pattern enforce the same invariant.
- Telemetry records both allowed and denied decisions with useful context.
- The architecture record and threat assumptions reflect the new control.
Remediate at the durable control point
Apply authorization at the domain or resolver layer that owns each field and mutation, propagate actor and tenant context cryptographically, cap cost based on backend work, isolate loaders and caches by security context, and restrict direct subgraph access. Generate policy tests from the schema so new fields cannot ship without an explicit decision.
Retest the invariant, not just the payload
Repeat the original operation with aliases, fragments and direct identifiers, then sample sibling fields or subgraphs using the same authorization helper. Confirm denied fields do not execute downstream work, legitimate partial responses remain correct and complexity controls behave consistently for allowed clients.
Use the NIST Technical Guide to Information Security Testing and Assessment for assessment planning and evidence discipline, and the OWASP Web Security Testing Guide for relevant application techniques. Adapt both to the system-specific trust decision described here.
The architecture decision
Scope GraphQL as an execution graph, not one URL. Confidence comes from testing field and mutation policy, traversal, cost, caching, federation and downstream effects across representative identities and tenants.
Ask WIMD to test the GraphQL execution graph from schema and resolvers through federation, caches and authoritative state.
