Category
Research
-
-
Why it’s critical to move beyond overly aggregated machine-learning metrics
-
At MIT, a continued commitment to understanding intelligence
-
3 Questions: How AI could optimize the power grid
-
Decoding the Arctic to predict winter weather
-
Japan Science and Technology Agency Develops NVIDIA-Powered Moonshot Robot for Elderly Care
-
MIT scientists investigate memorization risk in the age of clinical AI
-
Google’s year in review: 8 areas with research breakthroughs in 2025
-
Marine Biological Laboratory Explores Human Memory With AI and Virtual Reality
-
Guided learning lets “untrainable” neural networks realize their potential
-
A new way to increase the capabilities of large language models
-
A “scientific sandbox” lets researchers explore the evolution of vision systems
-
Evaluating AI’s ability to perform scientific research tasks
-
Measuring AI’s capability to accelerate biological research
-
How confessions can keep language models honest
-
Early experiments in accelerating science with GPT-5
-
How evals drive the next chapter in AI for businesses
-
Understanding neural networks through sparse circuits
-
Introducing IndQA
-
Defining and evaluating political bias in LLMs
-
Sora 2 is here
-
How people are using ChatGPT
-
Why language models hallucinate
-
Early methods for studying affective use and emotional well-being on ChatGPT
-
Introducing deep research
-
OpenAI o3-mini System Card
-
OpenAI o3-mini
-
Computer-Using Agent
-
Trading inference-time compute for adversarial robustness
-
OpenAI o1 System Card
-
Advancing red teaming with people and AI
-
Introducing SimpleQA
-
Simplifying, stabilizing, and scaling continuous-time consistency models
-
Evaluating fairness in ChatGPT
-
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
-
Learning to reason with LLMs
-
OpenAI o1-mini
-
Introducing SWE-bench Verified
-
Improving Model Safety Behavior with Rule-Based Rewards
-
GPT-4o mini: advancing cost-efficient intelligence
-
Prover-Verifier Games improve legibility of language model outputs
-
OpenAI and Los Alamos National Laboratory announce research partnership
-
Consistency Models
-
A Holistic Approach to Undesired Content Detection in the Real World
-
Improved Techniques for Training Consistency Models
-
Extracting Concepts from GPT-4
-
Hello GPT-4o
-
Understanding the source of what we see and hear online
-
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
-
Video generation models as world simulators
-
Building an early warning system for LLM-aided biological threat creation
-
Improving mathematical reasoning with process supervision
-
Democratic inputs to AI
-
GPTs are GPTs: An early look at the labor market impact potential of large language models
-
GPT-4
-
Point-E: A system for generating 3D point clouds from complex prompts
-
Scaling laws for reward model overoptimization
-
Introducing Whisper
-
Efficient training of language models to fill in the middle
-
DALL·E 2 pre-training mitigations
-
Learning to play Minecraft with Video PreTraining
-
Evolution through large models
-
Techniques for training large neural networks
-
Teaching models to express their uncertainty in words
-
Hierarchical text-conditional image generation with CLIP latents
-
A research agenda for assessing the economic impacts of code generation models
-
Solving (some) formal math olympiad problems
-
Text and code embeddings by contrastive pre-training
-
WebGPT: Improving the factual accuracy of language models through web browsing
-
Solving math word problems
-
TruthfulQA: Measuring how models mimic human falsehoods
-
Introducing Triton: Open-source GPU programming for neural networks
-
Evaluating large language models trained on code
-
Multimodal neurons in artificial neural networks
-
Understanding the capabilities, limitations, and societal impact of large language models
-
Scaling Kubernetes to 7,500 nodes
-
DALL·E: Creating images from text
-
CLIP: Connecting text and images
-
Generative language modeling for automated theorem proving
-
Image GPT
-
Language models are few-shot learners
-
AI and efficiency
-
Jukebox
-
Improving verifiability in AI development
-
OpenAI Microscope
-
Scaling laws for neural language models
-
Dota 2 with large scale deep reinforcement learning
-
Deep double descent
-
Procgen Benchmark
-
GPT-2: 1.5B release
-
Solving Rubik’s Cube with a robot hand
-
Emergent tool use from multi-agent interaction
-
GPT-2: 6-month follow-up
-
MuseNet
-
Generative modeling with sparse transformers
-
OpenAI Five defeats Dota 2 world champions
-
Implicit generation and generalization methods for energy-based models
-
Neural MMO: A massively multiagent game environment
-
Better language models and their implications
-
Computational limitations in robust classification and win-win results