In the swiftly evolving landscape of artificial intelligence, discussions about the future capabilities of large language models (LLMs) and the trajectory of their development are increasingly pivotal. This article analyses the positions of key thought leaders in AI regarding the scalability of these models and juxtaposes these views with historical trends like Moore’s law.
This analysis is intended for AI researchers, developers in the tech industry, and business leaders invested in AI advancements, as understanding these perspectives can significantly influence strategic planning and investment decisions.
Benedict Evans’ Position on AI Scaling
Benedict Evans suggests that unlike physical sciences, which have robust theoretical models to predict outcomes (like rocketry in physics), we lack a comprehensive theory of intelligence that can predict the development and capabilities of AGI. This absence of a solid theoretical framework makes it difficult to predict AI’s trajectory as precisely as we can in fields governed by established laws like physics.
In comparing AI’s progress to Moore’s Law, Benedict Evans expresses skepticism about applying such a straightforward predictive model to artificial intelligence development. He argues that while Moore’s Law provided a reliable framework for anticipating advancements in semiconductor technology based on historical data, the complexities and unknowns in AI, particularly in large language models, don’t allow for such clear-cut projections.
Moore’s law a good or bad analogy?

While recognizing the absence of a model of intelligence that offers precise, physics-like predictions, the analogy of Moore’s Law actually provides a valuable framework for understanding the potential trajectory of artificial intelligence. It’s important to emphasize that Moore’s Law is an observed trend rather than a strict scientific law. It reflects historical data on how the number of transistors on microchips has doubled approximately every two years, enhancing computing power without prescriptive insight into the mechanisms or limitations of such growth.

Not knowing HOW there will be more compute available didn’t impact people’s ability to reliably build for a future where that additional compute would be available, evident in the strategic planning of PC games in the mid-2000s with development cycles spanning over two years. In a similar vein, industry leaders from OpenAI, Anthropic, and Google DeepMind all indicate that the capability to predict performance increases through scaling is a significant factor driving the current push towards ever-larger models.

Sam Altman (OpenAI): Sam Altman emphasizes the continued improvement of AI models with increased computational power and data. He believes in aggressively pursuing scaling as a viable path to enhance AI capabilities, positing that larger models have historically shown improved performance, which justifies the investment in scaling up resources. He supports the notion that if the trend continues, AI will achieve significant breakthroughs:
Sholto Douglas (Anthropic): Sholto discusses the impact of scaling on the efficiency and capabilities of AI models. He notes that increasing compute resources does lead to noticeable improvements in model performance. Douglas underscores the importance of not just more resources but also better model architecture and training techniques to fully harness the potential of scaling:
Demis Hassabis (Google DeepMind): Demis articulates a cautious but optimistic view on scaling. He acknowledges that while scaling has led to improvements in AI capabilities, there are diminishing returns and potential barriers that may not be predictable by simply extrapolating from past trends. He stresses the importance of innovation in algorithms and architectures alongside scaling to sustain growth in AI capabilities:
Implications of AI Scaling Trends

Given the current trajectory and historical trends analogous to Moore’s Law, we can reasonably expect continued improvements in AI capabilities in the short term. This expectation is supported by the consistent past performance increases as computational power and model sizes have expanded, and the belief of those at the top that this trend is likely to continue.
While scaling up AI models has historically boosted capabilities, the expected quadratic relationship between model size and performance implies diminishing returns on capital investment. Despite this trend, the prospect of scaling models up to a hundred times larger than current versions remains viable without even factoring in refinements in architecture and improving efficiency to maximize the returns.
Sam advises companies to plan and build with an expectation of continuous improvements in AI capabilities. Businesses are encouraged to develop strategies that are adaptable to rapid advancements in AI technology:


Leave a Reply