dataset_id stringlengths 3 22 | pdf_id_x int64 0 7 | pdf_name stringclasses 8
values | pdf_title stringclasses 8
values | page_number int64 1 679 | base64_str stringlengths 24.5k 11.9M | base64_bytes int64 24.5k 11.9M | transcription stringlengths 45 7.45k ⌀ | transcription_input_tokens float64 0 850 | transcription_output_tokens float64 0 1.62k | pdf_id_y null | error null |
|---|---|---|---|---|---|---|---|---|---|---|---|
1_1 | 1 | 1.pdf | Journal of Geophysics | 1 | "iVBORw0KGgoAAAANSUhEUgAACbAAAA21CAIAAABqgr3/AAAACXBIWXMAAC4jAAAuIwF4pT92AAVvzklEQVR4nOzdB3wVVd7/8Vj(...TRUNCATED) | 475,176 | "Cc J 3} NIEDERSACHSISCHE STAATS- UND ~ L UNIVERSITATSBIBLIOTHEK GOTTINGEN Werk Jahr: 1980 Kollek(...TRUNCATED) | 850 | 277 | null | null |
1_2 | 1 | 1.pdf | Journal of Geophysics | 2 | "iVBORw0KGgoAAAANSUhEUgAACgYAAA3CCAIAAABmrx35AAAACXBIWXMAAC4jAAAuIwF4pT92AAdlOUlEQVR4nOzde0CUVf74cbB(...TRUNCATED) | 646,324 | "Journal Gy Geophysics JA SUSC AIA TLR Geophysik Volume 48 1980 Managing Editors W. Dieminger J.(...TRUNCATED) | 850 | 170 | null | null |
1_3 | 1 | 1.pdf | Journal of Geophysics | 3 | "iVBORw0KGgoAAAANSUhEUgAACMUAAAw6CAIAAADy/YFaAAAACXBIWXMAAC4jAAAuIwF4pT92AA4Vg0lEQVR4nOzdd7xV1Z0/7ms(...TRUNCATED) | 1,230,788 | "Journal of Geophysics — Zeitschrift fur Geophysik This journal was founded by the Deutsche Geoph(...TRUNCATED) | 850 | 608 | null | null |
1_4 | 1 | 1.pdf | Journal of Geophysics | 4 | "iVBORw0KGgoAAAANSUhEUgAACgYAAA3CCAIAAABmrx35AAAACXBIWXMAAC4jAAAuIwF4pT92ABTMYElEQVR4nOzdeXwV1f3/8QQ(...TRUNCATED) | 1,817,492 | "Author Index Alekseev A.S. 161 173 Baranskiy L.N. 1 Baumgardt D.R. 124 Baumjohann W. 7 Bock (...TRUNCATED) | 850 | 937 | null | null |
1_5 | 1 | 1.pdf | Journal of Geophysics | 5 | "iVBORw0KGgoAAAANSUhEUgAACgYAAA3CCAIAAABmrx35AAAACXBIWXMAAC4jAAAuIwF4pT92AA2xk0lEQVR4nOzdebxVVd0HfsB(...TRUNCATED) | 1,196,676 | "IV Inverse Seismic Problems A Characteristic Method for Numerical Solution of the Inverse Kinemat(...TRUNCATED) | 850 | 558 | null | null |
1_6 | 1 | 1.pdf | Journal of Geophysics | 6 | "iVBORw0KGgoAAAANSUhEUgAACMQAAAxmCAIAAABmcyLoAAAACXBIWXMAAC4jAAAuIwF4pT92AIfffUlEQVR4nNy9aZPjSLIkWP/(...TRUNCATED) | 11,872,868 | "Journal of | Geophysics Zettschrift fur Geophysik Volume 48 Number1 1980 ledersGchsische Staats- (...TRUNCATED) | 850 | 55 | null | null |
1_7 | 1 | 1.pdf | Journal of Geophysics | 7 | "iVBORw0KGgoAAAANSUhEUgAACOAAAAxECAIAAABqoSV/AAAACXBIWXMAAC4jAAAuIwF4pT92AAs2OklEQVR4nOy9vY4VSfK4fTQ(...TRUNCATED) | 979,808 | "Journal of Geophysics — Zeitschrift fur Geophysik Edited for the Deutsche Geophysikalische Gesel(...TRUNCATED) | 850 | 1,084 | null | null |
1_8 | 1 | 1.pdf | Journal of Geophysics | 8 | "iVBORw0KGgoAAAANSUhEUgAACMQAAAw6CAIAAAAdP+pkAAAACXBIWXMAAC4jAAAuIwF4pT92AA0HtElEQVR4nOzdfYxd5X0n8PG(...TRUNCATED) | 1,138,692 | "J. Geophys. 48 1-6 1980 Original Investigations Journal of Geophysics The Analysis of Simultane(...TRUNCATED) | 850 | 1,071 | null | null |
1_9 | 1 | 1.pdf | Journal of Geophysics | 9 | "iVBORw0KGgoAAAANSUhEUgAACMQAAAw5CAIAAACbq5jKAAAACXBIWXMAAC4jAAAuIwF4pT92ADPqx0lEQVR4nOy9ebzV4/r/f84(...TRUNCATED) | 4,536,692 | "Table 1. List of stations Station Nomen- Geo- Geo- L-para- Local clature magnetic magnetic meter (...TRUNCATED) | 850 | 925 | null | null |
1_10 | 1 | 1.pdf | Journal of Geophysics | 10 | "iVBORw0KGgoAAAANSUhEUgAACMQAAAw5CAIAAACbq5jKAAAACXBIWXMAAC4jAAAuIwF4pT92AC4GBUlEQVR4nOzdZ3RV1dr/fdJ(...TRUNCATED) | 4,021,700 | "18 Oct 1974 A Midnight 20F AQT LIN KNG T/A ry) HEL 1 G ; te ns 5 é : - pe \\\\e] {t u is \\\\g (...TRUNCATED) | 850 | 733 | null | null |
GDZ Scientific Document Retrieval Benchmark
A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.
Dataset Structure
The dataset is divided into two operational configurations:
1. queries
Contains the curated evaluation question set.
- Fields:
pdf_id,page_number,question,answer - Size: 180 needle-in-the-haystack questions.
Query Generation
The 180 needle-in-the-haystack questions were generated with Gemini Flash 3.5 using a prompt similar to IRPAPPERS designed to produce highly specific, self-contained, single-page-answerable questions. The model was shown each page image and instructed as follows:
I need your help simulating information seeking questions that information retrieval (IR) scientists would ask about the images attached to this message.
Could you please write a question that each page of the attached images uniquely answers? Further, these questions will be used to find this particular page amongst a large collection of other research papers, therefore, please try to make the question highly specific to this particular paper. Don't make the question too complicated, it should have a relatively short answer, not more than one or two sentences.
Please further explain why this question is only answerable by the cited page and not others, and why this question is only answerable by this particular paper and not other IR papers. Please also provide the answer to the question.
Please follow these instructions when generating these questions, answers, and explanations:
- Avoid phrases such as "this study" — make sure the questions are self-contained, as if searching through a large collection of similar papers for the answers. Use the specific name of the paper if needed, but generally avoid such references.
- Avoid "the researchers" — reference the study by the paper's name if needed.
- Do not reference the study as "the paper" — always use the paper's name. Questions must be self-contained.
- Avoid mathematical equations in answers.
- Keep answers short — no more than one or two sentences.
Each generated question was manually filtered for quality and page-uniqueness before inclusion in the final query set.
2. docs
Contains the complete target corpus text and image database.
- Fields:
dataset_id,pdf_id_x,pdf_name,pdf_title,page_number,base64_str,base64_bytes,transcription,transcription_input_tokens,transcription_output_tokens,pdf_id_y,error - Size: ~3,000 document pages.
- Format: Tab-separated value format compressed via Gzip (
.csv.gz), ensuring raw text alignment and completely unquoted ("") Base64 image strings.
How It Was Built
- Corpus Selection: Selected historical geophysics & mathematics volumes from the GDZ containing complex scientific text, tables, formulas, and diagrams (English + German).
- Visual Processing: Extracted high-resolution page images (300 DPI) and encoded them into raw Base64 layout feature matrices (
base64_str). - Text Extraction: Processed all pages through Tesseract OCR to build a parallel searchable text layer (
transcription). - Query Generation: Generated and manually filtered 180 highly specific verification questions targeting fine-grained methodological parameters and equations buried across the ~3,000 pages.
Usage
You can load either configuration directly into your evaluation pipeline using the Hugging Face datasets library:
from datasets import load_dataset
# 1. Load the 180 question evaluation set
queries_dataset = load_dataset("Trungdaik/Visual_information_retrieval", "queries", split="train")
print("Sample Query:", queries_dataset[0])
# 2. Load the ~3,000 target document pages corpus
docs_corpus = load_dataset("Trungdaik/Visual_information_retrieval", "docs", split="train")
print("Sample Target Page:", docs_corpus[0]["dataset_id"])
- Downloads last month
- 17