Advanced
Open
Pro
Shipping R1-Zero-Style Pure RLVR Directly to a Customer-Facing Chatbot
Your team has reproduced something like DeepSeek's R1-Zero result: RL directly against a base model with verifiable math/code rewards only, no SFT warm start, no preference-based stages. Internal eval shows strong accuracy on your benchmark. A PM proposes shipping this model directly to a customer support chatbot, to save the engineering time of building out the rest of the R1-style pipeline (SFT cold start, a further RL stage, rejection sampling, distillation).
- Name the specific, predictable problem class you'd expect in the shipped model's actual outputs, and explain precisely why RLVR-only training does not constrain against it.
- Of "SFT cold start" and "the second, broader RL stage across reasoning and general tasks," which is the more direct fix for the problem in part 1, and why doesn't more of the same RLVR training fix it on its own?
- Is skipping the rejection-sampling and distillation stages defensible for this use case specifically, distinct from your answer about the cold start? Justify separately.
Share this question