Artificial Intelligence Content Hub

Beyond Moore’s Law: The Full Stack Driving AI in 2026


For decades, Moore’s Law was the definitive, tried and true method for describing the rate of change in computing, from a technical perspective and on a cost and performance curve. The impact of that one curve changed and developed the modern world.

AI has no single equivalent, as it involves a much more complicated and interconnected set of variables, so it also requires a new lens for measuring raw underlying capability. For example, cheap tokens does not necessarily mean cheap AI or a good outcome. What matters is the cost of completing useful work at an acceptable quality.

This cost-per-qualified-task lens helps analyze the companies within the ROBO Global Artificial Intelligence Index (THNQ) that are helping move and shape those Pareto-optimal curves, from both the hardware and software infrastructure sides, alongside the closed and open-source sides of an argument now at the forefront of politics in Washington.

Key Takeaways:

  • The cost of AI inference at a constant level of model performance fell roughly tenfold a year between 2021 and 2024, a factor of 1,000 in three years, according to Andreessen Horowitz. MIT researchers put the more recent frontier rate at five to 10 times annually.
  • That rate is faster than transistor scaling ever delivered, and, unlike Moore’s Law, it is a composite of several technological curves compounding rather than one.
  • As cost per qualified task falls, more AI workflows become economically viable, expanding opportunity across infrastructure, software, security, and applications that range from autonomous vehicles to personal agents.

See more: AI News You Need to Know, June Edition: Capex, Inference, & Beyond

Why Moore’s Law Is No Longer Enough

Moore’s Law is not over, but it no longer describes the complete picture or whole system. Leading-edge chips continue to improve, while performance also comes from specialized architectures, memory, networking, and the software above them.

It does not move as one curve. Silicon improves in steady generational steps. Model architecture improves in jumps, when someone changes how the software works. A single blended rate averages across both describes neither.

The better lens is the price of a completed task that meets the required standard. Hardware affects how efficiently a model runs. Model design affects the computation and approach required. Serving software, data, security, and human review affect whether the result is usable, or scalable.

The Full-Stack Race

Last week, at AMD’s Advancing AI 2026 event, AMD and Cerebras announced plans for a new inference system that pairs AMD’s Helios platform with the Cerebras Wafer-Scale Engine in a single workflow. AMD (AMD) handles prompt processing and high-volume throughput, while Cerebras (CBRS) generates the answer. Company modeling projects up to five times more tokens per second per watt than a Cerebras-only configuration. The announcement shows how system architecture can add another efficiency curve.

Meanwhile, this week, China AI company Moonshot AI released Kimi K3, a new open-weight model that shows the software layer moving independently. It activates only 16 of 896 specialist components for each token it processes. Moonshot estimates that the design improves overall scaling efficiency by about 2.5 times over Kimi K2.

Going deeper into the application and deployment layers, Nvidia (NVDA) also launched the Open Secure AI Alliance this week, bringing together companies across cloud computing, cybersecurity, enterprise software, and AI research. The alliance treats identity, permissions, guardrails, logs, and evaluation as part of agent security. It extends the coalition strategy Nvidia began at the model layer with the Nemotron Coalition, which pools research, data, evaluations, and computation across AI labs.

From Token Prices to Qualified Outcomes

Cost per token alone does not tell an investor what useful work costs. A cheaper model may require more attempts or more human review. A more expensive model may complete the same task correctly in one pass.

Some tasks are deterministic: The code passes, the numbers reconcile, or the required field is present. Others span chains of dependencies, several parties, and real-world interactions. Non-deterministic means the same process may not produce the same result each time, and reasonable people may prefer different outcomes. The modern economy handles that ambiguity through standards, review, and accountability.

Artificial Analysis offers one of the best current views of the trade-off. Its index covers agentic work, coding, scientific reasoning, and general knowledge, and reports average cost per benchmark task. Claude Opus 5 scores 61 at $2.03 per task. DeepSeek V4 Flash scores 44 at four cents. At the end of the day, a lower-scoring model may be the economic choice when it meets the standard requirements, and thus, we see deflationary impact on the cost of “intelligence.”

The limitation is the word “task.” Benchmarks grade work inside a controlled environment. Valuable business processes often cross software systems, organizations, and the physical world, or depend on subjective judgment. Cost per task is a useful frontier, but not a universal price tag.

The industry is moving toward that frontier. OpenAI frames cost per successful task around price, compute, and the likelihood of reaching the right result. Nvidia uses “intelligence per dollar” for refining models after initial training. Both put capability and quality alongside cost.

How the Opportunity Expands

As the cost of a qualified outcome falls, more work becomes economical to automate or augment. Agents may consume more tokens as they plan, use tools, and recover from errors. This even expands into the physical realm, where agents can summon or interact on behalf of individuals and organizations in the real world, or even act as embodied agents in real life. This amplification of scope, with access to systems and data, increases cybersecurity requirements. Agents running autonomously require watching observability, at a scale that dwarfs previous monitoring needs. The opportunity runs from chips and cloud infrastructure to networking, security, data platforms, business processes, and industry-specific applications.

Nebius Group NBIS, an AI cloud provider, illustrates the enabling layer. Its Nvidia partnership spans AI-factory design, inference software, hardware deployment, and fleet management. It was also an early launch partner for Kimi K3, an example of the ecosystem benefiting when open-source models win.

THNQ captures several layers pushing this frontier and commercializing the resulting capabilities. Moore’s Law taught investors to watch one curve. AI requires watching how chips, system architecture, models, data, security, and deployment improve together. Companies lowering the cost of useful work, or expanding what AI can do, are building the next phase of the market.

THNQ is the underlying index for the ROBO Global Artificial Intelligence ETF (THNQ) and the L&G Artificial Intelligence UCITS ETF (AIAI.LN).

Looking for regular updates? Subscribe here for weekly insights on robotics, AI, and healthcare technology, delivered straight to your inbox.

For more news, information, and analysis, visit the Artificial Intelligence Content Hub.

VettaFi LLC (“VettaFi”) is the index provider for THNQ and ROBO, for which it receives an index licensing fee. However, these funds are not issued, sponsored, endorsed, or sold by VettaFi, and VettaFi has no obligation or liability in connection with the issuance, administration, marketing, or trading of these funds.

Loading...