Dataset Viewer
Duplicate
The dataset viewer is not available for this split.
Cannot load the dataset split (in streaming mode) to extract the first rows.
Error code:   StreamingRowsError
Exception:    CastError
Message:      Couldn't cast
lfs_expected: bool
media_type: string
path: string
sha256: string
size_bytes: int64
bundle: struct<filename: string, sha256: string, size_bytes: int64>
  child 0, filename: string
  child 1, sha256: string
  child 2, size_bytes: int64
current_compatible: bool
loadable: bool
schema_version: int64
remote_prefix: string
upload_complete: bool
archive_only: bool
status: string
inner_integrity: struct<checksums_entries: int64, checksums_sha256: string, inventory_records: int64, inventory_sha25 (... 10 chars omitted)
  child 0, checksums_entries: int64
  child 1, checksums_sha256: string
  child 2, inventory_records: int64
  child 3, inventory_sha256: string
payload_commit_oid: string
verification: struct<fresh_cache_download: bool, inner_checksums_verified: bool, lfs_pointer_rejected: bool, tar_s (... 21 chars omitted)
  child 0, fresh_cache_download: bool
  child 1, inner_checksums_verified: bool
  child 2, lfs_pointer_rejected: bool
  child 3, tar_sha256_verified: bool
to
{'archive_only': Value('bool'), 'bundle': {'filename': Value('string'), 'sha256': Value('string'), 'size_bytes': Value('int64')}, 'current_compatible': Value('bool'), 'inner_integrity': {'checksums_entries': Value('int64'), 'checksums_sha256': Value('string'), 'inventory_records': Value('int64'), 'inventory_sha256': Value('string')}, 'loadable': Value('bool'), 'payload_commit_oid': Value('string'), 'remote_prefix': Value('string'), 'schema_version': Value('int64'), 'status': Value('string'), 'upload_complete': Value('bool'), 'verification': {'fresh_cache_download': Value('bool'), 'inner_checksums_verified': Value('bool'), 'lfs_pointer_rejected': Value('bool'), 'tar_sha256_verified': Value('bool')}}
because column names don't match
Traceback:    Traceback (most recent call last):
                File "/src/services/worker/src/worker/utils.py", line 149, in get_rows_or_raise
                  return get_rows(
                      dataset=dataset,
                  ...<4 lines>...
                      column_names=column_names,
                  )
                File "/src/libs/libcommon/src/libcommon/utils.py", line 272, in decorator
                  return func(*args, **kwargs)
                File "/src/services/worker/src/worker/utils.py", line 129, in get_rows
                  rows_plus_one = list(itertools.islice(safe_iter(ds, dataset=dataset), rows_max_number + 1))
                File "/src/services/worker/src/worker/utils.py", line 489, in safe_iter
                  yield from ds.decode(False) if ds.features else ds
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2818, in __iter__
                  for key, example in ex_iterable:
                                      ^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2355, in __iter__
                  for key, pa_table in self._iter_arrow():
                                       ~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2380, in _iter_arrow
                  for key, pa_table in self.ex_iterable._iter_arrow():
                                       ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 536, in _iter_arrow
                  for key, pa_table in iterator:
                                       ^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 419, in _iter_arrow
                  for key, pa_table in self.generate_tables_fn(**gen_kwags):
                                       ~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 343, in _generate_tables
                  self._cast_table(pa_table, json_field_paths=json_field_paths),
                  ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
                  pa_table = table_cast(pa_table, self.info.features.arrow_schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2369, in table_cast
                  return cast_table_to_schema(table, schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2297, in cast_table_to_schema
                  raise CastError(
                  ...<3 lines>...
                  )
              datasets.table.CastError: Couldn't cast
              lfs_expected: bool
              media_type: string
              path: string
              sha256: string
              size_bytes: int64
              bundle: struct<filename: string, sha256: string, size_bytes: int64>
                child 0, filename: string
                child 1, sha256: string
                child 2, size_bytes: int64
              current_compatible: bool
              loadable: bool
              schema_version: int64
              remote_prefix: string
              upload_complete: bool
              archive_only: bool
              status: string
              inner_integrity: struct<checksums_entries: int64, checksums_sha256: string, inventory_records: int64, inventory_sha25 (... 10 chars omitted)
                child 0, checksums_entries: int64
                child 1, checksums_sha256: string
                child 2, inventory_records: int64
                child 3, inventory_sha256: string
              payload_commit_oid: string
              verification: struct<fresh_cache_download: bool, inner_checksums_verified: bool, lfs_pointer_rejected: bool, tar_s (... 21 chars omitted)
                child 0, fresh_cache_download: bool
                child 1, inner_checksums_verified: bool
                child 2, lfs_pointer_rejected: bool
                child 3, tar_sha256_verified: bool
              to
              {'archive_only': Value('bool'), 'bundle': {'filename': Value('string'), 'sha256': Value('string'), 'size_bytes': Value('int64')}, 'current_compatible': Value('bool'), 'inner_integrity': {'checksums_entries': Value('int64'), 'checksums_sha256': Value('string'), 'inventory_records': Value('int64'), 'inventory_sha256': Value('string')}, 'loadable': Value('bool'), 'payload_commit_oid': Value('string'), 'remote_prefix': Value('string'), 'schema_version': Value('int64'), 'status': Value('string'), 'upload_complete': Value('bool'), 'verification': {'fresh_cache_download': Value('bool'), 'inner_checksums_verified': Value('bool'), 'lfs_pointer_rejected': Value('bool'), 'tar_sha256_verified': Value('bool')}}
              because column names don't match

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

No dataset card yet

Downloads last month
46