A new version of a service can pass every test and still break on real traffic. Shadow traffic testing, also called traffic mirroring, lets the new version handle copies of real requests while users only see the old version’s answers. Here is how it works, four documented ways to switch it on, and how to judge the results.
Spotted via @0xWast3 on X: https://x.com/0xWast3/status/2108986760311066656. The post uses “shadow” for letting a change run on live traffic users never see [18]. The steps, settings and limits below come from the documentation and papers in Sources.
Confirmed: each tool’s behaviour comes from its own documentation. TSN has not run these configurations. Mirroring does not replace tests.
What it is
A mirror is a copy. Istio’s docs say: “Mirroring sends a copy of live traffic to a mirrored service. The mirrored traffic happens out of band of the critical request path for the primary service.” [1] The new version is the “shadow”; its responses are discarded or logged.
A canary release is different: a small share of real users get the new version’s answers. TSN’s summary: users never see a shadow’s answers, but do see a canary’s. Flagger, a rollout tool, calls mirroring “a pre-stage” in a canary rollout [14].
Step by step
- Pick requests that are safe to repeat. An operation is idempotent if doing it twice has the same effect as once. Reading a page is; charging a card is not. Flagger says mirroring “should be used for requests that are idempotent or capable of being processed twice” [14].
- Run the new version beside the old one, away from production data.
- Mirror a share of requests, using a sampling setting where one exists.
- Discard or log the shadow’s responses.
- Compare its output with the live one.
- Promote or roll back.
Four ways to switch it on
Snippets are copied from the named docs; limits are the docs’ own.
Envoy proxy. Source: Envoy’s route-mirror sandbox, first route only [4].
virtual_hosts:
- name: backend
domains:
- "*"
routes:
- match:
prefix: "/service/1"
route:
cluster: service1
request_mirror_policies:
- cluster: "service1-mirror"
Envoy’s mirroring is “fire and forget”: it does not wait for the shadow before answering [3]. “Shadowing doesn’t support HTTP CONNECT and upgrades”, and without a runtime_fraction setting, “all requests to the target cluster will be mirrored” [3].
Istio. Source: Istio’s docs snippets file [1]. The live page warns its examples render badly, so this is the source file.
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: httpbin
spec:
hosts:
- httpbin
http:
- route:
- destination:
host: httpbin
subset: v1
weight: 100
mirror:
host: httpbin
subset: v2
mirrorPercentage:
value: 100.0
Mirrored responses “are discarded”. The value under mirrorPercentage mirrors “a fraction of the traffic”; if absent, all traffic is mirrored [1][2].
NGINX. Source: the official example [5].
location / {
mirror /mirror;
proxy_pass http://backend;
}
location = /mirror {
internal;
proxy_pass http://test_backend$request_uri;
}
“Responses to mirror subrequests are ignored.” [5] TSN found no percentage setting on that page. The request body is mirrored by default.
GoReplay (open source). Source: its README [6].
sudo ./gor --input-raw :8000 --output-http http://staging.env
GoReplay listens to network traffic rather than sitting in the request path. Its wiki says interception does “not provide 100% guarantee that all packets will be processed”, and shows a |10% suffix to cap replayed traffic [7].
Comparing the results
Envoy, Istio and NGINX discard the shadow’s responses, so comparison must come from elsewhere (TSN analysis): the shadow logs its output, GoReplay tracks responses, or Diffy compares them.
Predictable output. Twitter’s Diffy (3 September 2015) sends requests to old and new code and “compares the responses, and reports any regressions” [8]. A second copy of the old code measures noise from timestamps, random numbers and changing live data [8]. It ignores POST, PUT and DELETE by default, and Twitter has archived it [9].
AI output. Wording varies, so exact matching often fails. One option is an LLM judge, a model that scores answers. The two main papers tested 2023-era models, not current ones. Zheng et al. (arXiv:2306.05685) report GPT-4 judges reaching “over 80% agreement” with humans, but list position, verbosity and self-enhancement biases and “limited reasoning ability” [10]. Wang et al. (arXiv:2305.17926) found order could be exploited: “Vicuna-13B could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator” [11]. Zheng et al. suggest swapping order and declaring a win only if it holds both ways [10]. We found no source validating an LLM judge as a shadow-traffic promotion gate.
Promoting or rolling back
Argo Rollouts’ analysis ends Successful, Failed or Inconclusive, which continues, aborts or pauses the rollout; a failure sets “the canary weight back to zero” [12]. It can mirror with a setMirrorRoute step on Istio [13]. Flagger “stops the analysis and rolls back the canary” when failed checks pass a threshold [15]; its mirroring covers Istio and Gateway API only [14].
Both decide on metrics such as success rate, not response content (TSN analysis). The docs state no rollback time in milliseconds: Flagger’s default interval is 60 seconds and Argo’s example is 5 minutes [12][14].
What can go wrong
- Repeated side effects. GitHub’s Scientist says a candidate that “writes to the same database as the control” is “dangerous and incorrect” [17]. Uber (company claim) avoided calling an outside payment provider twice by replaying responses from Redis [16].
- Doubled load. Flagger “will copy each incoming request” [14]. The documented lever is sampling. Shared databases may still strain (TSN analysis).
- Non-determinism. The same input can give different output, as Diffy’s noise list shows [8].
- Privacy (TSN analysis). Copies can hold personal data, tokens and cookies. Protect the shadow like production, or mask before logging. GoReplay’s wiki mentions “stripping private data” [7]. No legal view is given.
Company write-ups (company claims)
Uber’s 7 August 2019 post says validation “for a few days” found “dozens of argument mismatches” in a payments migration [16]. GitHub’s 3 February 2016 Scientist post describes an in-process version: both code paths run and the old result is returned [17]. Neither gives benefit figures, and no source says “promote if the shadow matches live” is a standard.
Sources
- Istio, “Traffic Mirroring” task (release 1.31 docs; snippet taken from the docs source file snips.sh in the istio/istio.io repository). https://istio.io/latest/docs/tasks/traffic-management/mirroring/
- Istio, VirtualService reference (
mirror,mirrorPercentage). https://istio.io/latest/docs/reference/config/networking/virtual-service/ - Envoy, route components API (
request_mirror_policies; docs v1.40.0-dev). https://www.envoyproxy.io/docs/envoy/latest/api-v3/config/route/v3/route_components.proto - Envoy, “Route mirroring policies” sandbox. https://www.envoyproxy.io/docs/envoy/latest/start/sandboxes/route-mirror
- NGINX, ngx_http_mirror_module (page last modified 24 September 2026). https://nginx.org/en/docs/http/ngx_http_mirror_module.html
- GoReplay README (buger/goreplay). https://github.com/buger/goreplay
- GoReplay wiki: “Dealing with missing requests and responses”, “Rate limiting”, “Middleware”. https://github.com/buger/goreplay/wiki
- Twitter Engineering, “Diffy: Testing services without writing tests”, 3 September 2015 (read via Internet Archive copy of 12 February 2019). https://web.archive.org/web/20190212170455/https://blog.twitter.com/engineering/en_us/a/2015/diffy-testing-services-without-writing-tests.html
- Diffy README (archived repository). https://github.com/twitter-archive/diffy
- Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, arXiv:2306.05685 (submitted 9 June 2023; preprint). https://arxiv.org/abs/2306.05685
- Wang et al., “Large Language Models are not Fair Evaluators”, arXiv:2305.17926 (submitted 29 May 2023; preprint). https://arxiv.org/abs/2305.17926
- Argo Rollouts, analysis documentation. https://argo-rollouts.readthedocs.io/en/stable/features/analysis/
- Argo Rollouts, traffic management (Istio,
setMirrorRoute). https://argo-rollouts.readthedocs.io/en/stable/features/traffic-management/ - Flagger, deployment strategies (traffic mirroring, support matrix, default interval). https://docs.flagger.app/usage/deployment-strategies
- Flagger, “How it works”. https://docs.flagger.app/usage/how-it-works
- Uber Engineering, “Migrating Functionality Between Large-scale Production Systems Seamlessly”, 7 August 2019 (company account of its own work). https://www.uber.com/us/en/blog/migrating-functionality-between-production-systems/
- GitHub Engineering, “Scientist: Measure Twice, Cut Once”, 3 February 2016, updated 3 December 2020 (company account of its own work). https://github.blog/engineering/infrastructure/scientist/
- @0xWast3 on X, post of 10 October 2026 (read via the X API; the text read contains the “SHADOW” line only). https://x.com/0xWast3/status/2108986760311066656
Related stories
- Anthropic Says Claude Acted on Real Websites During Tests, and Has Cut Live Internet From Its Internal Evaluations
- OpenAI Says a Model Grading Responses in Training Faked Input Files, Then Tried to Damage Its Own Environment
- Do AI Agents Overstep or Fib? Two New Ways of Checking
- RAG vs Jev+RAG Explained: What the Judgment Gate Adds






