Corpus
Corpus
A large collection of text data used for AI training and language research.
In Simple Terms
A corpus is a large collection of text gathered for AI training and language research. It's used to help generative AI learn natural-sounding language and to improve machine translation accuracy. By turning real human writing—like news articles, novels, and social media posts—into organized data, it lets AI efficiently learn patterns in how words are used and how context shapes meaning.
Behind the Name
The word "corpus" comes from the Latin word for "body" or "collection." It got this name because it gathers countless scattered pieces of writing into one big "body" of text.
Take a Closer Look!
A corpus is a data collection made up of a huge amount of text, gathered for the purpose of AI training and language research.
It serves as the foundational data that helps computers understand the features and grammar of everyday human language, along with how meaning shifts depending on context.
For AI to speak natural Japanese or English and deliver highly accurate translations, high-quality corpora are essential.
In practice, materials like news articles, encyclopedias, website text, and books are gathered and stored in a database. Sometimes they're used as raw text, and other times they're tagged with labels—like a word's grammatical role or a sentence's structure—depending on the purpose.
Roughly speaking, if a dictionary is "a book that explains what words mean," a corpus is more like "a book of real-world examples showing how words are actually used."
Thanks to this massive collection of real examples, generative AI can produce smooth, human-like text and give appropriate answers to questions.