JAEN
RESEARCH 03

Cultural Heritage & Ancient Documents

CULTURAL HERITAGE / ANCIENT DOCUMENTS

Carrying fading culture into the future with AI — recognizing, restoring, and preserving ancient documents and oracle bone inscriptions through digital technology.

Research Topics  ›  Cultural Heritage & Ancient Documents

背景と意義

Background

Japan holds a vast quantity of classical books and ancient documents that remain unorganized, while China's roughly 3,000-year-old oracle bone inscriptions are difficult to decipher due to aging and complex character forms. These historical materials are a shared intellectual asset of humanity, yet deciphering, organizing, and preserving them demands extensive expertise and time.

We use AI technologies centered on character recognition, image restoration, and document understanding to improve the readability, information extraction, and organization of historical materials. Our goal is to accelerate the digital preservation, research, and transmission of cultural heritage, contributing to both academic research and public access.

技術アプローチ

Methods

Oracle Bone Inscription Recognition

Oracle bone inscriptions are ancient Chinese characters from the Yin dynasty, difficult to recognize automatically due to their variety of forms and age-related degradation. We combine image processing and deep learning to detect, classify, and organize these characters, supporting the digitization and study of oracle bone materials.

Cursive Japanese (Kuzushiji) Recognition

Most classical Japanese books are written in cursive kuzushiji script, where connected characters and shifting forms make automatic reading difficult. We combine character-candidate extraction, character recognition, and document structure analysis to support reading and streamline the organization of classical books.

Restoration of Historical Document Images

Ancient documents and classical books suffer from degradation, stains, fading, and low contrast, all of which reduce readability. We use image restoration technologies including GANs and diffusion models to improve visibility while preserving textual information, supporting the preservation and use of historical materials.

Document Understanding & Knowledge Extraction

Beyond character recognition, we extract vocabulary, themes, and chronological information from materials to build semantic understanding of documents. This supports systematizing historical materials, improving searchability, and enabling knowledge discovery, opening new possibilities for cultural heritage research.

Anywhere, Anytime Recognition of Ancient Documents

Being able to recognize documents anywhere, anytime is an important way to make classical books easier to organize. We build a deep-learning recognition model on a server, combining a web-based system that transfers document images over the web for recognition with an Android-based system that runs on tablets and similar devices. Web-based recognition system ↗

代表的な研究成果

Selected Works