Dataset Viewer
Auto-converted to Parquet Duplicate
dataset_id
stringlengths
3
22
pdf_id_x
int64
0
7
pdf_name
stringclasses
8 values
pdf_title
stringclasses
8 values
page_number
int64
1
679
base64_str
stringlengths
24.5k
11.9M
base64_bytes
int64
24.5k
11.9M
transcription
stringlengths
45
7.45k
⌀
transcription_input_tokens
float64
0
850
transcription_output_tokens
float64
0
1.62k
pdf_id_y
null
error
null
1_1
1
1.pdf
Journal of Geophysics
1
"iVBORw0KGgoAAAANSUhEUgAACbAAAA21CAIAAABqgr3/AAAACXBIWXMAAC4jAAAuIwF4pT92AAVvzklEQVR4nOzdB3wVVd7/8Vj(...TRUNCATED)
475,176
"Cc J 3} NIEDERSACHSISCHE STAATS- UND ~ L UNIVERSITATSBIBLIOTHEK GOTTINGEN Werk Jahr: 1980 Kollek(...TRUNCATED)
850
277
null
null
1_2
1
1.pdf
Journal of Geophysics
2
"iVBORw0KGgoAAAANSUhEUgAACgYAAA3CCAIAAABmrx35AAAACXBIWXMAAC4jAAAuIwF4pT92AAdlOUlEQVR4nOzde0CUVf74cbB(...TRUNCATED)
646,324
"Journal Gy Geophysics JA SUSC AIA TLR Geophysik Volume 48 1980 Managing Editors W. Dieminger J.(...TRUNCATED)
850
170
null
null
1_3
1
1.pdf
Journal of Geophysics
3
"iVBORw0KGgoAAAANSUhEUgAACMUAAAw6CAIAAADy/YFaAAAACXBIWXMAAC4jAAAuIwF4pT92AA4Vg0lEQVR4nOzdd7xV1Z0/7ms(...TRUNCATED)
1,230,788
"Journal of Geophysics — Zeitschrift fur Geophysik This journal was founded by the Deutsche Geoph(...TRUNCATED)
850
608
null
null
1_4
1
1.pdf
Journal of Geophysics
4
"iVBORw0KGgoAAAANSUhEUgAACgYAAA3CCAIAAABmrx35AAAACXBIWXMAAC4jAAAuIwF4pT92ABTMYElEQVR4nOzdeXwV1f3/8QQ(...TRUNCATED)
1,817,492
"Author Index Alekseev A.S. 161 173 Baranskiy L.N. 1 Baumgardt D.R. 124 Baumjohann W. 7 Bock (...TRUNCATED)
850
937
null
null
1_5
1
1.pdf
Journal of Geophysics
5
"iVBORw0KGgoAAAANSUhEUgAACgYAAA3CCAIAAABmrx35AAAACXBIWXMAAC4jAAAuIwF4pT92AA2xk0lEQVR4nOzdebxVVd0HfsB(...TRUNCATED)
1,196,676
"IV Inverse Seismic Problems A Characteristic Method for Numerical Solution of the Inverse Kinemat(...TRUNCATED)
850
558
null
null
1_6
1
1.pdf
Journal of Geophysics
6
"iVBORw0KGgoAAAANSUhEUgAACMQAAAxmCAIAAABmcyLoAAAACXBIWXMAAC4jAAAuIwF4pT92AIfffUlEQVR4nNy9aZPjSLIkWP/(...TRUNCATED)
11,872,868
"Journal of | Geophysics Zettschrift fur Geophysik Volume 48 Number1 1980 ledersGchsische Staats- (...TRUNCATED)
850
55
null
null
1_7
1
1.pdf
Journal of Geophysics
7
"iVBORw0KGgoAAAANSUhEUgAACOAAAAxECAIAAABqoSV/AAAACXBIWXMAAC4jAAAuIwF4pT92AAs2OklEQVR4nOy9vY4VSfK4fTQ(...TRUNCATED)
979,808
"Journal of Geophysics — Zeitschrift fur Geophysik Edited for the Deutsche Geophysikalische Gesel(...TRUNCATED)
850
1,084
null
null
1_8
1
1.pdf
Journal of Geophysics
8
"iVBORw0KGgoAAAANSUhEUgAACMQAAAw6CAIAAAAdP+pkAAAACXBIWXMAAC4jAAAuIwF4pT92AA0HtElEQVR4nOzdfYxd5X0n8PG(...TRUNCATED)
1,138,692
"J. Geophys. 48 1-6 1980 Original Investigations Journal of Geophysics The Analysis of Simultane(...TRUNCATED)
850
1,071
null
null
1_9
1
1.pdf
Journal of Geophysics
9
"iVBORw0KGgoAAAANSUhEUgAACMQAAAw5CAIAAACbq5jKAAAACXBIWXMAAC4jAAAuIwF4pT92ADPqx0lEQVR4nOy9ebzV4/r/f84(...TRUNCATED)
4,536,692
"Table 1. List of stations Station Nomen- Geo- Geo- L-para- Local clature magnetic magnetic meter (...TRUNCATED)
850
925
null
null
1_10
1
1.pdf
Journal of Geophysics
10
"iVBORw0KGgoAAAANSUhEUgAACMQAAAw5CAIAAACbq5jKAAAACXBIWXMAAC4jAAAuIwF4pT92AC4GBUlEQVR4nOzdZ3RV1dr/fdJ(...TRUNCATED)
4,021,700
"18 Oct 1974 A Midnight 20F AQT LIN KNG T/A ry) HEL 1 G ; te ns 5 é : - pe \\\\e] {t u is \\\\g (...TRUNCATED)
850
733
null
null
End of preview. Expand in Data Studio

GDZ Scientific Document Retrieval Benchmark

A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.

Dataset Structure

The dataset is divided into two operational configurations:

1. queries

Contains the curated evaluation question set.

  • Fields: pdf_id, page_number, question, answer
  • Size: 180 needle-in-the-haystack questions.

Query Generation

The 180 needle-in-the-haystack questions were generated with Gemini Flash 3.5 using a prompt similar to IRPAPPERS designed to produce highly specific, self-contained, single-page-answerable questions. The model was shown each page image and instructed as follows:

I need your help simulating information seeking questions that information retrieval (IR) scientists would ask about the images attached to this message.

Could you please write a question that each page of the attached images uniquely answers? Further, these questions will be used to find this particular page amongst a large collection of other research papers, therefore, please try to make the question highly specific to this particular paper. Don't make the question too complicated, it should have a relatively short answer, not more than one or two sentences.

Please further explain why this question is only answerable by the cited page and not others, and why this question is only answerable by this particular paper and not other IR papers. Please also provide the answer to the question.

Please follow these instructions when generating these questions, answers, and explanations:

  • Avoid phrases such as "this study" — make sure the questions are self-contained, as if searching through a large collection of similar papers for the answers. Use the specific name of the paper if needed, but generally avoid such references.
  • Avoid "the researchers" — reference the study by the paper's name if needed.
  • Do not reference the study as "the paper" — always use the paper's name. Questions must be self-contained.
  • Avoid mathematical equations in answers.
  • Keep answers short — no more than one or two sentences.

Each generated question was manually filtered for quality and page-uniqueness before inclusion in the final query set.

2. docs

Contains the complete target corpus text and image database.

  • Fields: dataset_id, pdf_id_x, pdf_name, pdf_title, page_number, base64_str, base64_bytes, transcription, transcription_input_tokens, transcription_output_tokens, pdf_id_y, error
  • Size: ~3,000 document pages.
  • Format: Tab-separated value format compressed via Gzip (.csv.gz), ensuring raw text alignment and completely unquoted ("") Base64 image strings.

How It Was Built

  1. Corpus Selection: Selected historical geophysics & mathematics volumes from the GDZ containing complex scientific text, tables, formulas, and diagrams (English + German).
  2. Visual Processing: Extracted high-resolution page images (300 DPI) and encoded them into raw Base64 layout feature matrices (base64_str).
  3. Text Extraction: Processed all pages through Tesseract OCR to build a parallel searchable text layer (transcription).
  4. Query Generation: Generated and manually filtered 180 highly specific verification questions targeting fine-grained methodological parameters and equations buried across the ~3,000 pages.

Usage

You can load either configuration directly into your evaluation pipeline using the Hugging Face datasets library:

from datasets import load_dataset

# 1. Load the 180 question evaluation set
queries_dataset = load_dataset("Trungdaik/Visual_information_retrieval", "queries", split="train")
print("Sample Query:", queries_dataset[0])

# 2. Load the ~3,000 target document pages corpus
docs_corpus = load_dataset("Trungdaik/Visual_information_retrieval", "docs", split="train")
print("Sample Target Page:", docs_corpus[0]["dataset_id"])
Downloads last month
17