AppSec · Oct 1, 2026 · 7 min read

GraphQL batching: brute-force amplification the rate limit never sees

A single HTTP request carrying 50 login mutations is one request to the rate limiter and 50 to the resolver - brute-force at the HTTP layer is gone; abuse at the operation layer has barely started.

GraphQL batching lets a client pack multiple operations into one HTTP request. The spec-level intent is efficiency - fewer round trips, one network hop. The security consequence is a blind spot: every HTTP-layer rate limiter, WAF rule and brute-force threshold that counts requests counts one, while the resolver processes fifty operations behind that single count. The lock on the front door is open; the room it guards has fifty doors the counter never saw.

01How batching works and why it creates a blind spot

The GraphQL specification allows an array of operation objects at the root of the request body. Each entry is a complete, independent operation with its own query, variables and optional operation name. A GraphQL server that supports batching executes each entry in the array and returns an array of results - one result object per input, in order. The entire exchange is a single HTTP POST.

REST (login) GraphQL - unbatched POST /login x 1 POST /graphql x 1 rate limiter sees 1 request rate limiter sees 1 request resolver handles 1 attempt resolver handles 1 attempt GraphQL - batched (50 mutations in one body) POST /graphql x 1 rate limiter sees 1 request ──► PASS resolver iterates array ──► 50 login attempts evaluated response: [{...},{...}, ... x50] amplification factor = batch size; HTTP count = 1
A rate limiter at the HTTP layer counts requests, not operations. Batching multiplies operations per request without increasing the counted unit.

02The abuse patterns the HTTP counter misses

Authentication brute-force is the most direct case. A single batched POST carrying fifty login mutation entries, each with a different password candidate for the same username, registers as one request against an IP-based or session-based rate limit. An attacker who knows the burst budget - say, five attempts per second - can send a 100-entry batch every 20 seconds and try 18,000 candidates in five minutes while the counter never exceeds 15. The same amplification applies to OTP brute-force, account enumeration, and any mutation that enforces its own per-call quota without being aware it is being called fifty times in one HTTP body.

batch-login-probe.graphqlgraphql
# A confirming probe uses 3 entries - enough to demonstrate the pattern
# without hammering the target. A real attack would use hundreds.
[
  { "query": "mutation L($u:String!,$p:String!){login(username:$u,password:$p){token}}",
    "variables": { "u": "alice@example.com", "p": "apNonce62615533a" } },
  { "query": "mutation L($u:String!,$p:String!){login(username:$u,password:$p){token}}",
    "variables": { "u": "alice@example.com", "p": "apNonce62615533b" } },
  { "query": "mutation L($u:String!,$p:String!){login(username:$u,password:$p){token}}",
    "variables": { "u": "alice@example.com", "p": "apNonce62615533c" } }
]
# Confirming artifact: server returns [{...},{...},{...}] - three results
# in one response and no 429. HTTP counter shows exactly 1 request.

Beyond authentication, batching amplifies any per-operation budget: OTP delivery rate limits, password-reset token generation, account enumeration on the userExists query, and even read amplification - a batch of 200 user-profile queries is a server-side request amplification attack against the database, not a DoS against the HTTP layer. The common thread is that HTTP-level tooling cannot see the distinction.

03Confirming the gap safely

A safe confirming probe sends a small batch - three to five entries - of a sensitive mutation with benign, nonce-bearing credentials that cannot match any real account. The confirming artifact is the response structure: the server returns an array with one result per input entry, no 429 status, and no error indicating that batching was rejected. That triple fact - array response, no rate-limit error, batch count matches input count - proves the server accepts and executes batched mutations without per-operation counting. No real credential is tested, no account state is modified and no real brute-force loop runs during the probe.

Proven - HighGraphQL batch accepted without per-operation rate limit - POST /graphql
POST /graphql [{login(u=alice,p=nonce-a)},{login(u=alice,p=nonce-b)},{login(u=alice,p=nonce-c)}] ──► HTTP 200 [{"data":null,"errors":[...]},{"data":null,"errors":[...]},{"data":null,"errors":[...]}] No 429 returned. Rate-limit counter showed 1 request for 3 evaluated operations. Single batch of 3 nonce entries. No real credentials used. HTTP log: 1 POST.
Oracle: response is an array of length 3, no 429, no batch-rejected error. Rated high - HTTP-layer rate limiting can be multiplied by batch size; critical escalation if OTP or account-lockout mutations are reachable via the same pattern.

Escalation to critical applies when a rate-sensitive mutation - OTP verification, password reset, account lockout bypass - is reachable via the same endpoint and the batch probe confirms that endpoint also accepts arrays. The confirming probe for the escalated finding uses a batch against that mutation, still with nonce credentials that cannot succeed, and still counting the HTTP requests observed versus the operations evaluated. Two separate probes, two separate findings, each rated at the capability it confirmed.

04Defense: move the limit inside the operation layer

An HTTP-layer rate limiter counts what it can see. Closing the batching gap requires at least one control that operates at the operation level, not the request level. The three most effective controls are independent and should be combined: a per-request operation count cap (reject any array longer than N), an operation-level rate limiter that counts each entry in the batch toward the caller's budget, and query complexity accounting that sums the cost of all batch entries before any resolver runs. None of these requires disabling batching entirely; they constrain the amplification factor to a value the server can sustain.

ControlWhere it sitsWhat it stopsGap if absent
Batch size capGateway / middlewareLimits entries per array to N (e.g. 5)Unlimited amplification of any mutation
Per-operation rate limitResolver middlewareCounts each array entry toward the caller's quotaHTTP-layer quota bypassed entirely by batching
Query complexity budgetExecution engineSums cost across all batch entries before resolvingComplexity limit applies per-item, not per-batch
Disable batch for mutationsGateway ruleForbids array body on mutation endpointsBrute-force and amplification via mutation still open

Server-side implementations vary. Apollo Server, Yoga and Mercurius each have configuration paths for batch limits and complexity; Hasura exposes per-role operation limits that apply per entry. The confirming probe documents which server library handled the request - that determines which configuration path closes the gap. A WAF rule that rejects an array body at the HTTP layer is a second line of defense, not a substitute for operation-level controls inside the execution engine.

1
HTTP request logged by the rate limiter
50+
operations evaluated per batched POST
0
real credentials used in the confirming probe

The underlying lesson is the same one that applies across amplification classes: a control is only as strong as the layer it sits at. Rate limits at the HTTP tier are correct and necessary - they are just not sufficient for an endpoint that multiplies its work per request. Closing the gap means placing at least one limit inside the execution path, where the operation count is visible.

Evidence over heuristicsReachable beats foundMonitor → Block: rolling out a gate without slowing teamsIaC-grounded threat modelingAir-gapped AppSec with no phone-homeOne platform vs three tools
See it on your own app →