Qwen 3 4B RLM RLVR Collection LoRA adapters, full fine-tuned checkpoints, and SFT warmup models trained with RLVR in the recursive language model depth-1 harness. • 14 items • Updated 30 days ago
Running on CPU Upgrade 265 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens 📝 265 Visualize synthetic‑data experiments as an interactive bookshelf
Qwen 3 4B RLM RLVR Collection LoRA adapters, full fine-tuned checkpoints, and SFT warmup models trained with RLVR in the recursive language model depth-1 harness. • 14 items • Updated 30 days ago
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-depth2-recursive-r64-a128-lr1e-5-adapter Reinforcement Learning • Updated Jun 13 • 3
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-depth2-recursive-r64-a128-lr1e-5-adapter Reinforcement Learning • Updated Jun 13 • 3
Qwen 3 4B RLM RLVR Collection LoRA adapters, full fine-tuned checkpoints, and SFT warmup models trained with RLVR in the recursive language model depth-1 harness. • 14 items • Updated 30 days ago