
It's no secret that we are reaching the limits of traditional AI methods. This challenge has been highlighted by some of the brightest minds in AI, including Ilya Sutskever, who has been vocal about these constraints.
Training runs for large models can cost tens of millions of dollars due to the cost of chips. The costs have become so high that Anthropic has estimated that it could cost as much to examine Claude's features as it did to develop it in the first place. Companies like Amazon are spending billions to erect new AI data centers in an effort to keep up with compute demands. While the recent release of DeepSeek’s R1 model sparked the conversation around scaling perhaps not being all you need, R1 is still subject to the limitations inherent to traditional AI.
But maybe all of this isn't necessary. With a better foundational understanding of how AI works, we can approach AI model control and deployment in new ways that require a fraction of the energy and compute.
As the recipient of the Endowed Chair's Fellowship at the University of California San Diego, my work since undergrad and beyond singularly focused on solving this problem. In my first year of grad school, I was invited to give a presentation at ICLR that examined the behavior of simpler models in conditions designed to make them more represenative of deep neural network behavior, leading to collaborations with kindred spirits at Stanford.
IInstead of scaling by brute force, we broke down how LLMs actually learn and developed methods to understand and control them from a first-principles perspective. The same month we started, we won the inaugural PyTorch Conference Startup Showcase Award. It was remarkable to be recognized by the most widely used framework for training deep learning models while actively working to redefine the paradigm.
Deep learning models feature artificial neurons vaguely similar to ours, filtering data through them and then backpropagating information to learn features. By eschewing the inefficiencies and less theoretically justified parts of generative models, we create a path forward to the next generation of truly intelligent AI.
In the short time we’ve been building CTGT, we’ve continually seen growing evidence that this approach is breaking through the wall traditional approaches have hit. By embracing the bias towards simplicity seen in “The Bitter Lesson,” we can more intelligently scale models with practical applications. While we've seen incredible advancements with deep learning over the past decade, we now need to build the next evolution of AI beyond deep learning.
CTGT is doing just that. You can read about our new round of funding here.
We're grateful to Gradient, General Catalyst, Y Combinator, Liquid 2 and all the amazing angels who share our vision for the next generation of AI, including Francois Chollet, creator of Keras, Paul Graham, co-founder of Y Combinator, and other luminaries.
If you're interested in working at the forefront of intelligence, join us.