Everything from my experiments on training RLMs.
Lorenzo
lsteno
AI & ML interests
None yet
Organizations
models 13
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-depth2-recursive-r64-a128-lr1e-5-adapter
Reinforcement Learning • Updated • 3
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-FullFT-lr1e-5-depth1-v1
4B • Updated • 3
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-FullFT-lr5e-6-depth1-v1
Text Generation • 4B • Updated • 22
lsteno/qwen3-rlm-depth1-r64-a128-lr1e-5-s150-bal35f40v1-lora
Updated • 1
lsteno/qwen3-rlm-depth1-r64-a128-lr5e-7-s150-bal35f40v1-lora
Updated • 4
lsteno/qwen3-rlm-depth1-r16-a32-lr1e-4-s150-bal35f40v1-lora
Updated
lsteno/qwen3-rlm-depth1-r16-a32-lr1e-5-s150-bal35f40v1-lora
Updated • 1
lsteno/qwen3-rlm-depth1-r16-a32-lr5e-7-s150-bal35f40v1-lora
Updated
lsteno/qwen3-rlm-depth1-r4-a8-lr1e-4-s150-bal35f40v1-lora
Updated • 1
lsteno/Qwen3-4B-Instruct-2507-RLM-RL-depth1-r4-a8-lr1e-5-s150-lora
Updated • 2