-
ssurface/cot-dialect-qwen3-4b-instruct-sft-l1
Text Generation • Updated • 14 -
ssurface/cot-dialect-qwen3-4b-instruct-sft-l2
Text Generation • Updated • 11 -
ssurface/cot-dialect-qwen3-4b-instruct-sft-l3
Text Generation • Updated • 11 -
ssurface/cot-dialect-qwen3-4b-instruct-sft-l4
Text Generation • Updated • 18
🤝 Open to Collab
Frolov Anatolii
ssurface
·
AI & ML interests
None yet
Recent Activity
updated a collection 11 days ago
CoT Compression Dialects updated a collection 11 days ago
CoT Compression Dialects updated a collection 11 days ago
CoT Compression DialectsOrganizations
Sft-new
GRPO no length punshiment SFT GSM8K
GRPO SFT GSM8K
Hallucination & Agentic-Error Detection Datasets
Hallucination + agentic-error datasets from Toucan-1.5M, judged by gpt-oss-120b.
GRPO SFT-Length-Punishment-GDPO SFT GSM8K
GRPO abstract reward SFT GSM8K
Qwen3-4B CoT Compression Study
LoRA adapters trained for 5 progressively shorter chain-of-thought styles on GSM8K, plus the eval artifacts behind the Pareto curve.
CoT Compression Dialects
-
ssurface/cot-dialect-qwen3-4b-instruct-sft-l1
Text Generation • Updated • 14 -
ssurface/cot-dialect-qwen3-4b-instruct-sft-l2
Text Generation • Updated • 11 -
ssurface/cot-dialect-qwen3-4b-instruct-sft-l3
Text Generation • Updated • 11 -
ssurface/cot-dialect-qwen3-4b-instruct-sft-l4
Text Generation • Updated • 18
Hallucination & Agentic-Error Detection Datasets
Hallucination + agentic-error datasets from Toucan-1.5M, judged by gpt-oss-120b.
Sft-new
GRPO SFT-Length-Punishment-GDPO SFT GSM8K
GRPO no length punshiment SFT GSM8K
GRPO abstract reward SFT GSM8K
GRPO SFT GSM8K
Qwen3-4B CoT Compression Study
LoRA adapters trained for 5 progressively shorter chain-of-thought styles on GSM8K, plus the eval artifacts behind the Pareto curve.