Ilya Sutskever has spent the better part of two decades building machines that learn, and he has rarely explained himself in public. That gap between influence and commentary is precisely why his name resurfaces every time the direction of AI gets argued about. He co-authored AlexNet, the 2012 neural network that shattered image recognition and opened the deep learning era. He helped write the paper that made neural machine translation work. He was OpenAI’s chief scientist through GPT-1, GPT-2, GPT-3 and GPT-4. In May 2024, he left and started over.
Toronto, Hinton, and a network most people had written off
Sutskever was born in Russia in 1986 and grew up in Toronto after his family emigrated. He reached the University of Toronto at a time when neural networks were considered a dead end by most of the field. Geoffrey Hinton was one of the few researchers still pushing them.
The 2012 ImageNet competition settled the argument. AlexNet, built by Sutskever with Alex Krizhevsky and Hinton, cut the top-5 error rate to 15.3% while the nearest competitor managed 26.2%. That is not an incremental gain. It is a demolition, and within two years nearly every serious vision lab had switched to convolutional networks.
Two years later he co-authored Sequence to Sequence Learning with Neural Networks with Oriol Vinyals and Quoc Le at Google. The encoder-decoder design in that paper became the default approach to translation and set up a decade of language modelling work.
Google bought DNNresearch, the three-person company he had formed with Hinton and Krizhevsky, in 2013. Sutskever spent about a year at Google Brain before the next move.
The OpenAI years and the bet on scale
In December 2015 he became one of OpenAI’s founders, alongside Sam Altman, Greg Brockman, Elon Musk, Wojciech Zaremba, John Schulman and a handful of others. As chief scientist he acted as the lab’s internal compass on one question: how far does scale alone take you?
He was early with his answer. GPT-1 shipped in 2018 at a size that looks quaint now. GPT-2 followed in 2019 and was initially withheld from public release on misuse grounds. GPT-3 arrived in 2020 with 175 billion parameters and showed that a single general-purpose model could write code, answer questions and compose prose without task-specific training. GPT-4, released in March 2023, made it obvious the approach had commercial legs.
Not everything about the company’s origins was settled amicably. The founding arrangement between Musk and Altman eventually became public through Elon Musk’s lawsuit against Sam Altman, which exposed how much of OpenAI’s early identity was negotiated behind closed doors.
November 2023: the vote he still defends
On 17 November 2023, OpenAI’s board fired Sam Altman. Sutskever sat on that board and voted with the majority. Within days he had signed the staff letter demanding Altman’s return, and he later said he deeply regretted his part in the whole affair. The episode cost him his board seat and a good deal of goodwill inside the company.
He has not walked back the substance. Months later he stood by his role in Altman’s ouster, describing it as an attempt to protect the mission rather than a personal falling-out.
The crisis also redistributed power. Once the leadership settled, much of the operational authority inside OpenAI consolidated around Greg Brockman, a co-founder who had been central to the company since its earliest days.
Safe Superintelligence: one goal, no product
In June 2024, Sutskever surfaced with Safe Superintelligence Inc., founded with Daniel Gross and Daniel Levy. The pitch was unusual in its narrowness.
- One stated objective: safe superintelligence, nothing else.
- No product roadmap, no API, no published model weights.
- Two offices, Palo Alto and Tel Aviv, and a deliberately small team.
- Backers including Sequoia, Andreessen Horowitz, DST Global and Greenoaks.
The money showed up anyway. SSI raised $1 billion at a $5 billion valuation in September 2024. By April 2025, reports put a new round near $2 billion at a $32 billion valuation, led by Greenoaks. That is roughly a sixfold markup in seven months for a company with nothing to sell, no customers and no demo.
The appetite is not unique to SSI. Investors have been paying up for frontier-lab exposure for a while, and the AI hedge fund Situational Awareness sold its public portfolio while holding on to its Anthropic stake.
The contrast with the rest of the field is the whole point. Labs like Moonshot AI have released history’s largest open model and compete on users, benchmarks and downloads. SSI has published no weights, shipped no API and shown no demo.
What he believes about the next decade
Pre-training is running out of road
At NeurIPS in December 2024, Sutskever told the room that pre-training as the field practises it will end, because the supply of human text is finite. “We have but one internet,” he said, which is about as quotable as he gets.
The gains after that come from compute spent at inference: models that think for longer, check their own work, search, and generate training data for themselves. Reasoning, verification and self-play are the new scaling axes, and his prediction lines up with what has happened to model releases since.
Alignment as a design problem, not a filter
Superalignment, the team Sutskever co-led with Jan Leike from mid-2023, was meant to solve alignment for systems smarter than humans. It was disbanded after both men left, and Leike joined Anthropic. SSI’s founding premise is that alignment cannot be bolted onto a finished model as a late-stage patch.
He also says strange things sometimes. In February 2022 he tweeted that today’s large neural networks “might be slightly conscious.” He is not a person who posts carelessly, and the line has been argued over ever since.
Why SSI’s first move matters more than its next funding round
Sutskever’s career is a record of one bet paying off repeatedly: that general methods trained at scale beat hand-engineered ones, and that the field should follow the gradient rather than the intuition of the moment. AlexNet, sequence-to-sequence and the GPT line all trace back to that conviction.
The open question is whether a lab with no product can hold a $32 billion valuation long enough to produce something that justifies it. Frontier training runs cost extraordinary sums, and the researchers who would build safe superintelligence are the same people every other lab is trying to hire, often with equity packages only a revenue-generating company can fund. Whether SSI keeps its no-distractions posture through the next two years, or eventually ships something to keep the lights on, is the thing worth watching. Everything else about Sutskever can be read off the models he has already helped build.

