Datasets:
The dataset viewer is not available for this subset.
Exception: SplitsNotFoundError
Message: The split names could not be parsed from the dataset config.
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 290, in _generate_tables
pa_table = paj.read_json(
io.BytesIO(batch), read_options=paj.ReadOptions(block_size=block_size)
)
File "pyarrow/_json.pyx", line 364, in pyarrow._json.read_json
File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
raise convert_status(status)
pyarrow.lib.ArrowInvalid: JSON parse error: Column() changed from object to string in row 0
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 286, in get_dataset_config_info
for split_generator in builder._split_generators(
~~~~~~~~~~~~~~~~~~~~~~~~~^
StreamingDownloadManager(base_path=builder.base_path, download_config=download_config)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 101, in _split_generators
pa_table = next(iter(self._generate_tables(**splits[0].gen_kwargs, allow_full_read=False)))[1]
~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 304, in _generate_tables
batch = json_encode_fields_in_json_lines(original_batch, json_field_paths)
File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 111, in json_encode_fields_in_json_lines
examples = [ujson_loads(line) for line in original_batch.splitlines()]
~~~~~~~~~~~^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 20, in ujson_loads
return pd.io.json.ujson_loads(*args, **kwargs)
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^
ValueError: Expected object or value
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/split_names.py", line 68, in compute_split_names_from_streaming_response
for split in get_dataset_split_names(
~~~~~~~~~~~~~~~~~~~~~~~^
path=dataset,
^^^^^^^^^^^^^
config_name=config,
^^^^^^^^^^^^^^^^^^^
token=hf_token,
^^^^^^^^^^^^^^^
)
^
File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 340, in get_dataset_split_names
info = get_dataset_config_info(
path,
...<6 lines>...
**config_kwargs,
)
File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 291, in get_dataset_config_info
raise SplitsNotFoundError("The split names could not be parsed from the dataset config.") from err
datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config.Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
YAML Metadata Warning:The task_categories "image-retrieval" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other
MAT-VB: MAthematical Text-Vision Benchmark
π Best Resource Paper Award β 2025 ACM/IEEE Joint Conference on Digital Libraries (JCDL)
Official Hugging Face repository for MAT-VB, a multimodal benchmark constructed from Wikipedia mathematical content designed to evaluate mathematical image captioning and text-to-image retrieval models.
π Links
- GitHub Repository: AIIRLab/MATVB
- Paper (IEEE Xplore): MAT-VB: MAthematical Text-Vision Benchmark
π Dataset Files Overview
| File Name | Format | Description |
|---|---|---|
images.zip |
Archive (.zip) |
Contains all ~4,000 Wikipedia math images (216 MB uncompressed). |
MATVB_train.json |
JSON | Training annotations. Maps image_id (key) to its ground-truth caption (value). |
MATVB_query_nl.json |
JSON | Natural language query set for text-to-image search (human-authored intent queries). |
MATVB_query_caption.json |
JSON | Caption query set for text-to-image search (original Wikipedia captions). |
MATVB_qrel.qrel |
TREC format | Evaluation relevance file mapping query IDs to target image IDs. |
π― Task Definitions & Splits
1. Mathematical Image Captioning
- Training Set: Use all entries in
MATVB_train.json(3,600 image-caption pairs). - Test Set: Any image in
images.zipwhose ID is not present inMATVB_train.jsonserves as a test sample (400 test images). Models generate captions for these test images during evaluation.
2. Text-to-Image Retrieval
The benchmark provides 400 test evaluation queries under two modalities sharing identical query IDs:
- Natural Language Queries (
MATVB_query_nl.json): Human-authored search queries representing intent-based user searches. - Caption Queries (
MATVB_query_caption.json): Queries built using original, detailed Wikipedia captions. - Ground Truth (
MATVB_qrel.qrel): Use standard TREC evaluation tools (e.g.,trec_eval) to score retrieval runs against this relevance file.
π How to Load and Use
You can download and parse the dataset files directly using huggingface_hub and standard Python libraries:
import json
import zipfile
from huggingface_hub import hf_hub_download
# 1. Download dataset files
repo_id = "AIIRLab/MATVB"
train_path = hf_hub_download(repo_id=repo_id, filename="MATVB_train.json", repo_type="dataset")
query_nl_path = hf_hub_download(repo_id=repo_id, filename="MATVB_query_nl.json", repo_type="dataset")
query_cap_path = hf_hub_download(repo_id=repo_id, filename="MATVB_query_caption.json", repo_type="dataset")
qrel_path = hf_hub_download(repo_id=repo_id, filename="MATVB_qrel.qrel", repo_type="dataset")
images_zip_path = hf_hub_download(repo_id=repo_id, filename="images.zip", repo_type="dataset")
# 2. Extract image archive
with zipfile.ZipFile(images_zip_path, 'r') as zip_ref:
zip_ref.extractall("./images")
# 3. Load JSON annotations
with open(train_path, "r") as f:
train_data = json.load(f) # dict: {image_id: caption}
with open(query_nl_path, "r") as f:
queries_nl = json.load(f) # dict: {query_id: query_text}
with open(query_cap_path, "r") as f:
queries_caption = json.load(f) # dict: {query_id: query_text}
print(f"Loaded {len(train_data)} training samples.")
print(f"Loaded {len(queries_nl)} NL retrieval queries.")
- Downloads last month
- 150