Wikiwand AI

Draft:Talkie

Talkie, a vintage language model From Wikipedia, the free encyclopedia

Talkie is a small language model developed by Nick Levine, David Duvenaud, and Alec Radford. It was announced in April 2026 and described by the developers as a vintage language model. Talkie is trained solely on pre-1931 texts that are in the Public Domain,[1] in order to reduce legal issues and liability with releasing model data.[2] The model family consists of a 13 billion parameter model called talkie-1930-13b-base and a post-trained checkpoint designed to power a chat interface, called talkie-1930-13b-it.[3]



Original authorsNick Levine, David Duvenaud, Alec Radford
ReleaseApril 2026
Available inEnglish
Quick facts Talkie, Original authors ...
Talkie
Original authorsNick Levine, David Duvenaud, Alec Radford
ReleaseApril 2026
Available inEnglish
LicenseApache license
Websitehttps://talkie-lm.com/
Close

Development

The initial idea was to build a model trained on historical data that could be used to explore whether models can forecast future events: what is predictable, and how far out events can be predicted.[2] The model was also developed in order to study cultural change, and model self-conception.[4][3] Another goal expressed by authors is testing whether language models can arrive at inventions or scientific discoveries. The authors cite a thought experiment proposed by Demis Hassabis, who asked whether a model trained on data up to 1911 could independently discover Albert Einstein's General Relativity theory.[5]

The term vintage language model is attributed to Owain Evans and describes language models trained only on historical text. The purpose of such models is to simulate language use from the past, and to study behavior of models not contaminated by contemporary content. Other vintage models include Ranke 4B, Mr Chatterbox or Machina Mirabilis. Talkie is also inspired by Calcifer Computing’s work on Temporal Language Models, able to represent temporal trends in language.[6]

The model was trained on 260 billion tokens of pre-1931 English text from sources like the Institutional Data Initiative, Common Pile or the Internet Archive. The 31 December 1930 cutoff is based on copyright term rules in the United States, where works published between 1923-1977 are protected for 95 years. The data includes books, newspapers, periodicals, scientific journals, patents, and case law. The developers experimented with various optical character recognition (OCR) methods and developed a dedicated vintage OCR system. The compute needed to train the model was provided by Anthropic.[7]

One of the main challenges with building vintage model is contamination of the model by anachronistic data from beyond the cutoff data. This is typically due to incorrect metadata, or editing notes added to the text. Because of this, talkie is for example aware of Franklin Delano Roosevelt and Adolf Hitler.[2]

A dedicated post-training pipeline was developed in order to fine tune a chatbot based on the base model. For this purpose, only historical structured texts, such as etiquette manuals, letter-writing manuals, cookbooks, dictionaries, and encyclopedias, were used, to avoid contamination of the model. Nevertheless, due to the fact that the model was also fine-tuned through synthetic chats with a Claude Opus model, some anachronisms were introduced.[1] Simon Willinson describes the base talkie model as a "vegan" model – one trained entirely on licensed or out-of-copyright data.[8] The chat model does not quality, because of the above mentioned contamination during fine tuning, when proprietary synthetic data generated by Claude Opus was introduced.

Reception

References

Related Articles

Timelines

Top Qs

Fact Checks