mradermacher/qwen2.5-coder-3b-instruct-spider-sft-grpo-GGUF Reinforcement Learning • 3B • Updated 2 days ago • 152