Feature Extraction
Transformers
Safetensors
AuriStream
audio
speech
language-model
auristream
custom_code
Instructions to use TuKoResearch/AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TuKoResearch/AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="TuKoResearch/AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TuKoResearch/AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit
AuriStream is a speech language model by Greta Tuckute and Klemen Kotar.
This model predicts cochlear tokens from a tokenizer such as WavCochCausalV8192.
Native training step-zero initialization for the 7B-Deep 40-prediction model. This exactly uses origin seed 1110 and historical source commit b5619758b6e74a124bf792f9b4e454154121f6df, matching the start of W&B run 724216un. The weights are untrained FP32 values produced before XLA/FSDP wrapping and are the seed-1110 artifact used in the origin-seed forensic comparison.
Model Details
| Parameter | Value |
|---|---|
| Parameters | ~8.41B |
| Layers | 96 |
| Hidden Size | 2560 |
| Attention Heads | 32 |
| Vocab Size | 8192 |
| Prediction Steps | 40 |
Usage
from transformers import AutoModel, AutoConfig
# Load with trust_remote_code for custom model
model = AutoModel.from_pretrained(
"TuKoResearch/AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit",
trust_remote_code=True,
)
# Or load config first
config = AutoConfig.from_pretrained("TuKoResearch/AuriStream7BDeep_40Pred_BigAudioDataset_500k-randinit", trust_remote_code=True)
Base Model Code
This checkpoint uses shared model code from TuKoResearch/AuriStream-base.
Tokenizer
This model uses cochlear tokens from WavCochCausalV8192.
- Downloads last month
- 61