Sakana AI Releases Karamaru, Edo-Period Classical Japanese Chatbot Trained on 25 Million Characters
Sakana AI
Sakana AI built a chatbot, Karamaru, that answers in Edo-period classical Japanese instead of modern text. It trained on 25 million characters of old books, so replies actually think like Edo Japan, not just sound old.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI has released Karamaru, a chatbot that doesn't just sprinkle archaic word endings onto modern Japanese sentences. Ask it something in plain contemporary Japanese and it answers as if it were actually a resident of Edo-period Japan, drawing on the worldview and vocabulary of that era rather than faking the accent. That distinction matters more than it sounds. Anyone can tell a large language model to write like it's from centuries ago, but Sakana AI wanted the substance of the reply, not just the flavor, to carry that period's assumptions and knowledge.
The approach behind Karamaru is continued pretraining, not retrieval. The team took ELYZA's existing Japanese-language model, Llama-3-ELYZA-JP-8B, and kept training it on roughly 25 million characters pulled from Edo-era books. That's a modest dataset by large language model standards, which is exactly why continued pretraining on top of an already-strong Japanese base model made more sense than starting from scratch.
Building that 25-million-character dataset took work from three separate academic efforts. The crowd-sourced transcription platform Minna de Honkoku contributed roughly 12 million characters drawn from 2,901 transcribed items, mostly Edo-period books along with things like woodblock-printed news sheets and letters. The National Institute of Japanese Literature's decade-long project to digitize historical texts made around 300,000 classical works available as digital images in the first place. And the ROIS-DS Center for Open Data in the Humanities supplied about a million characters of human-transcribed text, plus roughly 12 million more characters that Sakana AI's own team pulled from 1,001 Edo-period books using an AI model called RURI, built for reading cursive kuzushiji script, then cleaned up with a correction tool they call OCR Refiner. Add it up and you get roughly 13 million human-transcribed characters and 12 million AI-transcribed ones.
Sakana AI argues this beats the usual method for historical-knowledge chatbots, which is retrieval-augmented generation: pulling matching text from an archive and having a modern model paraphrase it. That method struggles to find relevant passages for every possible question, and even top models like GPT-4o, when simply prompted to answer in Edo-period Japanese, tend to just tack archaic-sounding word endings onto otherwise modern content. Karamaru's continued pretraining, by contrast, lets both the content and the phrasing shift together.
The name itself is a small piece of trivia worth keeping. It nods to Tsutaya Jūzaburō, the Edo-period publisher who went by the pen name Tsutano Karamaru, and doubles as a pun on the tangled web of words and concepts any language model learns to weave. This release, Llama-3-Karamaru-v1, is explicitly a first version, posted on HuggingFace with a public demo for research and educational use, and Sakana AI says larger and more varied training data could follow in future iterations.
My take — AI-written commentary, not fact-checked reporting
Reviving centuries-old texts through a chatbot instead of a museum plaque is a genuinely clever use of AI, and one that doesn't require pretending the model is flawless. The dataset here is tiny compared to what usually trains these systems, and that's the real story: you don't always need billions of tokens to make a model useful, you need the right tokens for the job. Historical language communities and other low-resource domains should be taking notes rather than waiting for someone to throw more compute at English.
Read more about this at: Sakana AI
Related stories
Sakana AI Launches Sakana Namazu, a Japanese-Specialized LLM API
Sakana AI ·
16
From Japan, Products the World Will Use: An Interview with Sakana AI's Head of Product Development
Sakana AI ·
48
Sakana AI Develops Proprietary Technology for SNS Visualization and Countering False and Misleading Information in Ministry of Internal Affairs and Communications Project
Sakana AI ·
11