Dataset Viewer
The dataset viewer is not available for this subset.
Cannot get the split names for the config 'default' of the dataset.
Exception:    SplitsNotFoundError
Message:      The split names could not be parsed from the dataset config.
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 290, in _generate_tables
                  pa_table = paj.read_json(
                      io.BytesIO(batch), read_options=paj.ReadOptions(block_size=block_size)
                  )
                File "pyarrow/_json.pyx", line 364, in pyarrow._json.read_json
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: JSON parse error: Column() changed from object to string in row 0
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 286, in get_dataset_config_info
                  for split_generator in builder._split_generators(
                                         ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      StreamingDownloadManager(base_path=builder.base_path, download_config=download_config)
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 101, in _split_generators
                  pa_table = next(iter(self._generate_tables(**splits[0].gen_kwargs, allow_full_read=False)))[1]
                             ~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 304, in _generate_tables
                  batch = json_encode_fields_in_json_lines(original_batch, json_field_paths)
                File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 111, in json_encode_fields_in_json_lines
                  examples = [ujson_loads(line) for line in original_batch.splitlines()]
                              ~~~~~~~~~~~^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 20, in ujson_loads
                  return pd.io.json.ujson_loads(*args, **kwargs)
                         ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^
              ValueError: Expected object or value
              
              The above exception was the direct cause of the following exception:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/split_names.py", line 68, in compute_split_names_from_streaming_response
                  for split in get_dataset_split_names(
                               ~~~~~~~~~~~~~~~~~~~~~~~^
                      path=dataset,
                      ^^^^^^^^^^^^^
                      config_name=config,
                      ^^^^^^^^^^^^^^^^^^^
                      token=hf_token,
                      ^^^^^^^^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 340, in get_dataset_split_names
                  info = get_dataset_config_info(
                      path,
                  ...<6 lines>...
                      **config_kwargs,
                  )
                File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 291, in get_dataset_config_info
                  raise SplitsNotFoundError("The split names could not be parsed from the dataset config.") from err
              datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config.

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

YAML Metadata Warning:The task_categories "image-retrieval" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other

MAT-VB: MAthematical Text-Vision Benchmark

πŸ† Best Resource Paper Award β€” 2025 ACM/IEEE Joint Conference on Digital Libraries (JCDL)

Official Hugging Face repository for MAT-VB, a multimodal benchmark constructed from Wikipedia mathematical content designed to evaluate mathematical image captioning and text-to-image retrieval models.


πŸ”— Links


πŸ“ Dataset Files Overview

File Name Format Description
images.zip Archive (.zip) Contains all ~4,000 Wikipedia math images (216 MB uncompressed).
MATVB_train.json JSON Training annotations. Maps image_id (key) to its ground-truth caption (value).
MATVB_query_nl.json JSON Natural language query set for text-to-image search (human-authored intent queries).
MATVB_query_caption.json JSON Caption query set for text-to-image search (original Wikipedia captions).
MATVB_qrel.qrel TREC format Evaluation relevance file mapping query IDs to target image IDs.

🎯 Task Definitions & Splits

1. Mathematical Image Captioning

  • Training Set: Use all entries in MATVB_train.json (3,600 image-caption pairs).
  • Test Set: Any image in images.zip whose ID is not present in MATVB_train.json serves as a test sample (400 test images). Models generate captions for these test images during evaluation.

2. Text-to-Image Retrieval

The benchmark provides 400 test evaluation queries under two modalities sharing identical query IDs:

  • Natural Language Queries (MATVB_query_nl.json): Human-authored search queries representing intent-based user searches.
  • Caption Queries (MATVB_query_caption.json): Queries built using original, detailed Wikipedia captions.
  • Ground Truth (MATVB_qrel.qrel): Use standard TREC evaluation tools (e.g., trec_eval) to score retrieval runs against this relevance file.

πŸš€ How to Load and Use

You can download and parse the dataset files directly using huggingface_hub and standard Python libraries:

import json
import zipfile
from huggingface_hub import hf_hub_download

# 1. Download dataset files
repo_id = "AIIRLab/MATVB"

train_path = hf_hub_download(repo_id=repo_id, filename="MATVB_train.json", repo_type="dataset")
query_nl_path = hf_hub_download(repo_id=repo_id, filename="MATVB_query_nl.json", repo_type="dataset")
query_cap_path = hf_hub_download(repo_id=repo_id, filename="MATVB_query_caption.json", repo_type="dataset")
qrel_path = hf_hub_download(repo_id=repo_id, filename="MATVB_qrel.qrel", repo_type="dataset")
images_zip_path = hf_hub_download(repo_id=repo_id, filename="images.zip", repo_type="dataset")

# 2. Extract image archive
with zipfile.ZipFile(images_zip_path, 'r') as zip_ref:
    zip_ref.extractall("./images")

# 3. Load JSON annotations
with open(train_path, "r") as f:
    train_data = json.load(f)  # dict: {image_id: caption}

with open(query_nl_path, "r") as f:
    queries_nl = json.load(f)  # dict: {query_id: query_text}

with open(query_cap_path, "r") as f:
    queries_caption = json.load(f)  # dict: {query_id: query_text}

print(f"Loaded {len(train_data)} training samples.")
print(f"Loaded {len(queries_nl)} NL retrieval queries.")
Downloads last month
150