LLM-powered Biographies
Eugene Yan
An author tested multiple large language models including GPT-4, GPT-3.5, and Claude to generate biographies of themselves, comparing accuracy and performance across different models. GPT-4 performed best with correct general themes but factual errors in education and career details, while older models like text-davinci-003 produced entirely fabricated information such as claiming the subject was a Canadian venture capitalist. The experiment demonstrates current limitations in LLM training data coverage and memorization patterns, with models showing varying degrees of hallucination regardless of their overall capability level.
Why it matters
Asking LLMs to generate biographies to get a sense of how they memorize and regurgitate.