BW Thumbnail Burgundy Noisy Data_Aug 2026.jpg
AIA Labs: The Future of Investment Intelligence

Noisy Data Breaks RLVR

Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs. However, creating the high-quality verifiable answers needed to train the model is labor-intensive and expensive, as highlighted by the boom of data labeling companies, such as Mercor. This creates a fundamental question for post-training: To what extent can RLVR, with algorithmic improvements, tolerate large-volume, noisy data?

Recent literature promotes a counter-intuitive idea: RLVR is robust to noisy data. Prior work shows that training on “100% incorrect” answers leads to only 5% lower performance than clean data and achieves higher performance than forward rewards (enclosing any answer in \boxed{}). This suggests we can throw cheap, messy annotations at an LLM, tweak the RLVR loss, and still significantly improve reasoning.

We find that this hypothesis is false. The “surprisingly” high effectiveness of RLVR on “100%” noisy data is due to the contamination in the synthetic noise — the claimed “100%” noisy data contains a significant portion of clean data. In this blog, we develop a more rigorous noisy data curation pipeline and show that noisy data is destructive to RLVR, impacting the test accuracy by over 9%. Even current advanced algorithmic improvements fail to mitigate the severe impact.

Paper: https://arxiv.org/abs/2603.16140

Repo: https://github.com/uiuc-kang-lab/rlvr-noisy-data

BW-AIA Blog_Noisy Data_01.png
Training on noisy data leads to significantly degraded performance compared to training on clean data.

Prior “Noisy” Data Contains Correct Labels

BW-AIA Blog_Noisy Data_02.png
Examples of correct labels contained in the prior “noisy” data.

Recent work claimed that LLMs could learn effectively from 100% incorrect annotations during RLVR. But when we verified the synthetic noisy datasets used in prior work, we found a critical flaw: the noise was contaminated with correct answers. This hidden contamination inflated the performance of models trained on it.

BW-AIA Blog_Noisy Data_03.png
Data re-verification pipeline for synthesizing a truly noisy dataset.

How did this happen? Prior work generated noisy labels for math datasets by sampling answers from a base LLM and filtering out the correct ones using basic symbolic equivalence checkers. However, this filtering fails for two primary reasons:

  1. Insufficient gold answers: A math problem often has multiple valid solutions, but ground-truth annotations typically only capture one (as shown in Example 1 above).
  2. Inadequate equivalence checking: Basic symbolic verifiers frequently fail to recognize when a generated answer is mathematically identical to the gold answer but formatted differently (as shown in Example 2 above).

To measure the actual impact of noise, we must ensure the noisy data has incorrect answers. We applied a more rigorous re-verification pipeline using GPT-5 Pro combined with manual expert review. We discovered that 16% of the supposedly "incorrect" labels in the original dataset were actually correct. By purging these leaked answers, we constructed a truly noisy dataset.

The Collapse of Reasoning

When we purged those hidden correct annotations, the performance collapsed. In our experiments using Qwen2.5-Math-7B, training on 100% truly incorrect annotations dropped average accuracy on five math benchmarks (AIME’24, AIME’25, AMC’23, AMC’24, and MATH500) by 9% compared to training on clean data. This performance was even worse than training with format-only rewards.

BW-AIA Blog_Noisy Data_04.png
Training on noisy data leads to significantly degraded performance compared to training on clean data.

We also found that noisy training data leads to weaker reasoning. When scaling the number of attempts (k > 1), pass@k dropped below the base model's baseline capabilities. Furthermore, noise penalizes deep exploration. Models trained on noisy data produced reasoning chains that were 5–24% shorter than those trained on clean data.

BW-AIA Blog_Noisy Data_05.png
Noise leads to weaker reasoning (lower Pass@k and shorter response length).

Existing Algorithmic Improvements Fail on Noisy Data

Prior research has proposed a wide range of algorithmic improvements to the vanilla GRPO. Can these existing algorithm improvements make RLVR more robust to noisy data? Unfortunately, our empirical results show that they cannot.

We considered state-of-the-art RLVR variants: SAPO, DAPO, TIS, DR. GRPO, and PGFC (an algorithm explicitly designed to mitigate the impact of noise). We tested them under a 50% noise rate, a reasonable rate identified on real-world text-to-SQL datasets. Unfortunately, all of these algorithms lead to more than 5.0% accuracy degradation on at least one benchmark compared to training with clean data. On AIME and AMC benchmarks, these algorithms achieve comparable or lower performance than training with only format rewards.

BW-AIA Blog_Noisy Data_06.png
Under 50% noise, all of the improved algorithms lead to > 5% accuracy degradation on at least one benchmark compared to training with clean data.

Conclusion

Our empirical evidence shows that the impact of noisy data is destructive and fundamental to RLVR. Algorithmic tweaks simply cannot yet compensate for performance loss due to noisy training data. If we want to improve LLM reasoning, high-quality data remains essential.

We released our validated noisy dataset, model checkpoints, and training scripts at: https://github.com/uiuc-kang-lab/rlvr-noisy-data. For more technical details, please refer to our paper: https://arxiv.org/abs/2603.16140.


This research paper is prepared by and is the property of Bridgewater Associates, LP and is circulated for informational and educational purposes only. There is no consideration given to the specific investment needs, objectives, or tolerances of any of the recipients. Additionally, Bridgewater’s actual investment positions may, and often will, vary from its conclusions discussed herein based on any number of factors, such as client investment restrictions, portfolio rebalancing and transactions costs, among others. Recipients should consult their own advisors, including tax advisors, before making any investment decision. This material is for informational and educational purposes only and is not an offer to sell or the solicitation of an offer to buy the securities or other instruments mentioned. Any such offering will be made pursuant to a definitive offering memorandum. This material does not constitute a personal recommendation or take into account the particular investment objectives, financial situations, or needs of individual investors which are necessary considerations before making any investment decision. Investors should consider whether any advice or recommendation in this research is suitable for their particular circumstances and, where appropriate, seek professional advice, including legal, tax, accounting, investment, or other advice. No discussion with respect to specific companies should be considered a recommendation to purchase or sell any particular investment. The companies discussed should not be taken to represent holdings in any Bridgewater strategy. It should not be assumed that any of the companies discussed were or will be profitable, or that recommendations made in the future will be profitable.

The information provided herein is not intended to provide a sufficient basis on which to make an investment decision and investment decisions should not be based on simulated, hypothetical, or illustrative information that have inherent limitations. Unlike an actual performance record simulated or hypothetical results do not represent actual trading or the actual costs of management and may have under or overcompensated for the impact of certain market risk factors. Bridgewater makes no representation that any account will or is likely to achieve returns similar to those shown. The price and value of the investments referred to in this research and the income therefrom may fluctuate. Every investment involves risk and in volatile or uncertain market conditions, significant variations in the value or return on that investment may occur. Investments in hedge funds are complex, speculative and carry a high degree of risk, including the risk of a complete loss of an investor’s entire investment. Past performance is not a guide to future performance, future returns are not guaranteed, and a complete loss of original capital may occur. Certain transactions, including those involving leverage, futures, options, and other derivatives, give rise to substantial risk and are not suitable for all investors. Fluctuations in exchange rates could have material adverse effects on the value or price of, or income derived from, certain investments.

Bridgewater research utilizes data and information from public, private, and internal sources, including data from actual Bridgewater trades. Sources include AERIC INC, BCA, Bloomberg Finance L.P., Candeal, Carbon Arc, CEIC Data Company Ltd., Ceras Analytics, China Bull Research, Citibank, Clarus Financial Technology, CLS Processing Solutions, Consensus Economics Inc., Consumer Edge, CRU Group, DTCC Data Repository, Ecoanalitica, Energy Aspects Corp, Enverus, EPFR Global, Eurasia Group, Evercore ISI, FactSet Research Systems, The Financial Times Limited, Finaeon, Inc., FINRA, GaveKal Research Ltd., GlobalSource Partners, Goldman Sachs, Harvard Business Review, Haver Analytics, Inc., IEA, Institutional Shareholder Services (ISS), The Investment Funds Institute of Canada, ICE Derived Data (UK), Investment Company Institute, International Institute of Finance, JP Morgan, JTSA Advisors, LSEG Data and Analytics, MarketAxess, Metals Focus Ltd, MSCI, Inc., National Bureau of Economic Research, Neudata, Organisation for Economic Cooperation and Development, Pensions & Investments Research Center, Pitchbook, Political Alpha, Renaissance Capital Research, Rhodium Group, RP Data, Rubinson Research, Rystad Energy, S&P Global Market Intelligence, Sentix GmbH, SGH Macro, Shanghai Metals Market, Smart Insider Ltd., Swaps Monitor, Tradeweb, United Nations, US Department of Commerce, Visible Alpha, Wells Bay, Wind Financial Information LLC, With Intelligence, Wood Mackenzie Limited, World Bureau of Metal Statistics, World Economic Forum, and YieldBook. While we consider information from external sources to be reliable, we do not assume responsibility for its accuracy. Data leveraged from third-party providers, related to financial and non-financial characteristics, may not be accurate or complete. The data and factors that Bridgewater considers within its research process may change over time.

This information is not directed at or intended for distribution to or use by any person or entity located in any jurisdiction where such distribution, publication, availability, or use would be contrary to applicable law or regulation, or which would subject Bridgewater to any registration or licensing requirements within such jurisdiction. No part of this material may be (i) copied, photocopied, or duplicated in any form by any means or (ii) redistributed without the prior written consent of Bridgewater® Associates, LP.

The views expressed herein are solely those of Bridgewater as of the date of this report and are subject to change without notice. Bridgewater may have a significant financial interest in one or more of the positions and/or securities or derivatives discussed. Those responsible for preparing this report receive compensation based upon various factors, including, among other things, the quality of their work and firm revenues.

Connecting the Dots
Sign up to receive insights and analysis from Bridgewater Associates
You're almost finished.
You will receive an email confirmation shortly.
There's been an error. Please start over and try again.
Connecting the Dots
Sign up to receive insights and analysis from Bridgewater Associates
This website uses cookies. Click here for additional details. By continuing to use this website, you consent to the use of cookies.

Internet Explorer is not supported by this website.

For optimal browsing we recommend using Chrome, Safari, or Firefox.