Re-export the 8da4w text variants at the sweep's group size
#1
by msluszniak - opened
The published 8da4w files predate the quantization sweep: they carry no quantized embedding table, which is why they are ~176 MB larger than these.
- 350m: group_size 32, embedding_quantize 8,32
- 1_2b: group_size 32, embedding_quantize 8,0
Also fixes the export emitting bf16 logits: the published artifacts and every config declare float32, and the quantized branch was overriding the dtype to bf16, so any re-export silently changed the output contract.
Configs gain get_n_layers (16), which was null in all four. The fp16 and MLX variants are untouched. Both .pte were probed against their config and match.
msluszniak changed pull request status to merged