Dataset Viewer
Duplicate
The dataset viewer is not available for this split.
Cannot load the dataset split (in streaming mode) to extract the first rows.
Error code:   StreamingRowsError
Exception:    CastError
Message:      Couldn't cast
version: string
data: struct<paragraphs: list<item: struct<context: string, qas: list<item: struct<answers: list<item: str (... 110 chars omitted)
  child 0, paragraphs: list<item: struct<context: string, qas: list<item: struct<answers: list<item: struct<answer_start: i (... 75 chars omitted)
      child 0, item: struct<context: string, qas: list<item: struct<answers: list<item: struct<answer_start: int64, text: (... 63 chars omitted)
          child 0, context: string
          child 1, qas: list<item: struct<answers: list<item: struct<answer_start: int64, text: string>>, id: string, is_imp (... 33 chars omitted)
              child 0, item: struct<answers: list<item: struct<answer_start: int64, text: string>>, id: string, is_impossible: bo (... 21 chars omitted)
                  child 0, answers: list<item: struct<answer_start: int64, text: string>>
                      child 0, item: struct<answer_start: int64, text: string>
                          child 0, answer_start: int64
                          child 1, text: string
                  child 1, id: string
                  child 2, is_impossible: bool
                  child 3, question: string
  child 1, title: string
epoch: double
train_loss: double
train_steps_per_second: double
train_runtime: double
total_flos: double
train_samples_per_second: double
-- schema metadata --
pandas: '{"index_columns": [], "column_indexes": [], "columns": [{"name":' + 323
to
{'epoch': Value('float64'), 'total_flos': Value('float64'), 'train_loss': Value('float64'), 'train_runtime': Value('float64'), 'train_samples_per_second': Value('float64'), 'train_steps_per_second': Value('float64')}
because column names don't match
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 290, in _generate_tables
                  pa_table = paj.read_json(
                      io.BytesIO(batch), read_options=paj.ReadOptions(block_size=block_size)
                  )
                File "pyarrow/_json.pyx", line 342, in pyarrow._json.read_json
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                  return check_status(status)
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: JSON parse error: Invalid value. in row 0
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/utils.py", line 147, in get_rows_or_raise
                  return get_rows(
                      dataset=dataset,
                  ...<4 lines>...
                      column_names=column_names,
                  )
                File "/src/libs/libcommon/src/libcommon/utils.py", line 272, in decorator
                  return func(*args, **kwargs)
                File "/src/services/worker/src/worker/utils.py", line 127, in get_rows
                  rows_plus_one = list(itertools.islice(safe_iter(ds, dataset=dataset), rows_max_number + 1))
                File "/src/services/worker/src/worker/utils.py", line 483, in safe_iter
                  yield from ds.decode(False) if ds.features else ds
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2840, in __iter__
                  for key, example in ex_iterable:
                                      ^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2373, in __iter__
                  for key, pa_table in self._iter_arrow():
                                       ~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2398, in _iter_arrow
                  for key, pa_table in self.ex_iterable._iter_arrow():
                                       ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 536, in _iter_arrow
                  for key, pa_table in iterator:
                                       ^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 419, in _iter_arrow
                  for key, pa_table in self.generate_tables_fn(**gen_kwags):
                                       ~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 339, in _generate_tables
                  yield Key(shard_idx, 0), self._cast_table(pa_table)
                                           ~~~~~~~~~~~~~~~~^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
                  pa_table = table_cast(pa_table, self.info.features.arrow_schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
                  return cast_table_to_schema(table, schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
                  raise CastError(
                  ...<3 lines>...
                  )
              datasets.table.CastError: Couldn't cast
              version: string
              data: struct<paragraphs: list<item: struct<context: string, qas: list<item: struct<answers: list<item: str (... 110 chars omitted)
                child 0, paragraphs: list<item: struct<context: string, qas: list<item: struct<answers: list<item: struct<answer_start: i (... 75 chars omitted)
                    child 0, item: struct<context: string, qas: list<item: struct<answers: list<item: struct<answer_start: int64, text: (... 63 chars omitted)
                        child 0, context: string
                        child 1, qas: list<item: struct<answers: list<item: struct<answer_start: int64, text: string>>, id: string, is_imp (... 33 chars omitted)
                            child 0, item: struct<answers: list<item: struct<answer_start: int64, text: string>>, id: string, is_impossible: bo (... 21 chars omitted)
                                child 0, answers: list<item: struct<answer_start: int64, text: string>>
                                    child 0, item: struct<answer_start: int64, text: string>
                                        child 0, answer_start: int64
                                        child 1, text: string
                                child 1, id: string
                                child 2, is_impossible: bool
                                child 3, question: string
                child 1, title: string
              epoch: double
              train_loss: double
              train_steps_per_second: double
              train_runtime: double
              total_flos: double
              train_samples_per_second: double
              -- schema metadata --
              pandas: '{"index_columns": [], "column_indexes": [], "columns": [{"name":' + 323
              to
              {'epoch': Value('float64'), 'total_flos': Value('float64'), 'train_loss': Value('float64'), 'train_runtime': Value('float64'), 'train_samples_per_second': Value('float64'), 'train_steps_per_second': Value('float64')}
              because column names don't match

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

No dataset card yet

Downloads last month
69