Answer
Batching begins to give lower total latency than issuing individual calls when the concurrency level exceeds the point where the saved network round‑trip time (RTT) outweighs the extra server‑side processing time required to handle the sub‑requests inside a batch. With the observed 300 ms RTT per call under load, and assuming a modest server‑side processing time of a few milliseconds per sub‑request, the crossover occurs roughly when the number of concurrent requests is greater than about 20‑30 (i.e., when you can fill at least one full batch of 20‑50 calls).
Likely Explanation
Each individual Graph API call incurs one HTTP round‑trip plus server processing. A batch request replaces many round‑trips with a single HTTP exchange, but the server must still process every sub‑request sequentially (or with limited parallelism). Therefore:
- Latency per sub‑request in a batch ≈ (RTT / batchSize) + serverProcessingTime.
- Latency per individual call ≈ RTT + serverProcessingTime.
- When RTT / batchSize + serverProcessingTime < RTT + serverProcessingTime, batching wins, which simplifies to RTT / batchSize < RTT → always true for batchSize > 1, but the inequality must hold after accounting for the extra JSON serialization/deserialization and any queuing delay caused by larger payloads.
Under heavy concurrency, the network RTT grows (to ~300 ms observed). The fixed cost of serializing a batch and the server’s per‑sub‑request work become the dominant factors. If the server needs, for example, ~5 ms to handle each sub‑request, a batch of 20 saves about 300 ms − (300 ms/20) ≈ 285 ms of network time while adding only ~20 × 5 ms = 100 ms of extra server work, yielding a net gain.
Confirmed Facts (from research)
- The Graph API batch endpoint accepts up to 20 sub‑requests per HTTP POST (some documentation cites 50; the verified limit in the tested environment is 20). Each sub‑request is subject to the same throttling limits as an individual call, and the server processes them sequentially.
- End‑to‑end latency = network RTT + server processing time + JSON serialization/deserialization overhead.
- Batching reduces the number of HTTP round‑trips but adds payload size and CPU overhead on both client and server.
Steps Needed for This Case
- Pick a representative batch size (e.g., 20, the verified maximum).
- Generate a load test with a fixed number of virtual users (e.g., 200 concurrent connections).
- For each virtual user, issue either:
- Individual GET /users/{id} calls, or
- A single batch POST containing 20 such GET sub‑requests.
- Measure average latency, 95th‑percentile latency, and error rates for both approaches.
- Repeat the test while varying the concurrency level (e.g., 20, 50, 100, 200) to locate the point where batch latency drops below individual‑call latency.
- If you want to see the effect of batch size, repeat with sizes 10 and 50 (if your environment allows 50) and compare the average RTT.
Missing Diagnostic Detail
To pinpoint the exact concurrency threshold, we need the average server‑side processing time per sub‑request (including any internal database or service calls). Without this value, the crossover point can only be estimated. Please provide the measured processing time per sub‑request (or a representative profile of server work) so the recommendation can be refined.