Comparison of Core Abilities of Four Large Language Models: GPT-3, GPT-4, LLaMA 2, and PaLM 2
Abstract
. Different large language models vary dramatically in their resource requirements, with some costing millions to train while others can run on a laptop. This paper examines four models that represent different design philosophies: GPT-3, GPT-4, LLaMA 2, and PaLM 2. GPT-3, with its 175 billion parameters, first showed that scaling up could unlock few-shot learning. GPT-4 went further by adding image understanding, though OpenAI never published the architecture. Meta took the opposite approach with LLaMA 2, releasing weights publicly so anyone could experiment. Google's PaLM 2 focused on languages beyond English, training on over 100 languages. After reviewing how each model works and where it performs best, the paper discusses what problems remain unsolved, including that models still make up facts, training still costs too much, and no one fully understands why certain prompts fail.