Sebastian Raschka@rasbtOct 11, 2026, 05:17@BiploveYadav The model (training) should work fine without format reward, it’s just an optional stylistic thing. If the model never gets the answers right then you’d need to start with a simpler dataset.打开原帖#511482