GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF Text Generation • 1B • Updated 10 days ago • 121k • 157
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF Text Generation • 1B • Updated 10 days ago • 241k • 297
Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning Paper • 2606.18831 • Published Jun 17 • 7