HomeAIReSI Preprint: An Automated Safety Loop Cuts X-Teaming Attack Success From 86.01%...

ReSI Preprint: An Automated Safety Loop Cuts X-Teaming Attack Success From 86.01% to 31.45%, the Authors Report

A preprint posted to arXiv on 8 October 2026 describes ReSI (“Recursive Safety Improvement”). It repeatedly attacks an AI model, retrains it and checks the result. The authors report it cut one attack test’s average success across four models from 86.01% to 31.45% [1].

Preprint. Not peer reviewed. The findings are the authors’ own, not independently replicated, and come from the authors’ own benchmark setup. It is by Jingnan Zheng, Dongcheng Zhang, Yi Zhang and 10 other authors (13 in all), under cs.CR [1].

How ReSI works

According to the abstract, in each round ReSI “applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model” [1].

Red-teaming means attacking a model on purpose to find ways to make it give unsafe answers. A Pareto gate on capability retention is a rule that rejects updates that make the model worse at its normal tasks. Section 2.4 describes it as requiring safety gains while preserving benign compliance and instruction following [1].

The reported result

Attack success rate is the share of attack attempts that get an unsafe response; lower is better. X-Teaming is a named attack test. The paper says it “exploits adaptive multi-turn interactions” (Section 3), meaning the attacker adapts over several turns of conversation, and cites a 2025 paper on multi-turn jailbreaks [1][2].

The abstract says ReSI “reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna’s leading frontier result of 56.69%” [1]. That sets an average of four models against one frontier model’s score. The abstract calls the four only “dense and mixture-of-experts models”, and says ReSI does this “largely preserving general capabilities” [1].

What this does not show

  • Not replicated. No outside team has repeated it; it is not peer reviewed.
  • The authors’ chosen attacks. The X-Teaming test covers 159 behaviours, scored by a model the authors chose, GPT-4o (Appendix B.1, B.2.2). The paper itself says its experiments “cover a limited set of models and configurations over a finite number of update rounds” (Section 6) [1].
  • Lower attack success is not the same as safe. Untested attacks may still work, and the abstract does not say this works on products.
  • Capability retention is the authors’ claim. The abstract says “largely”. The paper says “Ministral shows larger losses on the reasoning and knowledge benchmarks” (Section 4.3), and that “Single-run comparisons do not establish statistical significance” (Appendix B.2.1) [1].

The Bottom Line

A preprint reports that an automated attack-and-retrain loop cut mean X-Teaming attack success across four models from 86.01% to 31.45%. It is the authors’ own, unreplicated, unreviewed result: a lower score on one set of attacks, not proof these models resist jailbreaks.

Related on TSN: Do AI Agents Overstep or Fib? Two New Ways of Checking

Sources

  1. Jingnan Zheng et al., “ReSI: Recursive Safety Improvement toward Resistant and Resilient AI”, arXiv:2610.12233 [cs.CR], v1, 8 October 2026 (preprint, primary source; HTML and PDF also checked). https://arxiv.org/abs/2610.12233
  2. Salman Rahman et al., “X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents”, arXiv:2504.13203, 15 April 2025 (cited by the ReSI paper). https://arxiv.org/abs/2504.13203

Share this story

More in this category

Latest on TSN

Free TSN tools: crypto calculator, Flux dashboard and more.