Thoughts on the Frontier Lab Economic Index
August 31, 2026
I wrote this in Spring 2025, when labs and funders were first discussing how to measure AI’s economic effects across labs.
Summary
Existing data on economic usage of LLMs is based only on consumer chat. Enterprise+API uses ~2x the tokens of consumer chat, and those tokens are skewed towards the research questions we’re interested in. Consumer-level data from OpenAI, Microsoft, Google, etc. will not help with the questions on firms, sectors, and industries.
But strategically, it’s better to aim for a mega-AEI (multi-lab Anthropic Economic Index) to get “a foot in the door”; I describe a barebones MVP below. Framing this as an evaluation could be really effective—”[This new model], by revealed preference of our user, is used x% more in augmenting productive work in sectors a, b, and c.”
Maybe thinking of it as an “index” is a bad idea; we don’t want a single number which can be gamed. We want a dataset, like the O*NET task data released by Anthropic, which can be interpreted by experts into a narrative.
There are two ways to think about this: research-focused and strategic. I have some thoughts on both, but I’m not sure what Schmidt’s strategic goals are here. So discount my thoughts on strategy accordingly.
What do researchers need and want?
Understanding what type of firms do and don’t adopt AI tools will be important. Those of us whom you are funding are studying what AI can do, not necessarily what it is doing. We are experimentalists, injecting AI into “business as usual” workplaces and measuring the effects. Observational data on who uses AI (well) is important to understanding takeoff speed.
Slow adoption can come from regulation, litigation risk, personal preference, energy constraints, etc. These are hard to detect in experimental data.
Consumer-level data from OpenAI, Microsoft, Google, etc. will not help with the questions on firms, sectors, and industries. For this, you need data on corporate clients. Enterprise+API uses ~2x the tokens of consumer chat, and those tokens are skewed towards the research questions we’re interested in: how AI is being used in the workplace, how industries will change, how jobs are being automated.
What can we do with an MVP?
Assuming we can’t get any corporate data, what could we learn from a mega-Anthropic Economic Index (mega-AEI)?
- Power would be helpful: putting in 100MM conversations instead of 1MM will enable detection of smaller effect sizes and more granular occupation-level analysis
- Cross-platform representation: People select into different platforms (e.g., my grandma only uses Gemini; doctors mostly use ChatGPT). Since the AEI data dropped, people have been assuming that Anthropic users are more sophisticated than OpenAI subscribers. General assumption is that the AEI sample overrepresents software developers. So with multiple labs’ data, we get a better cross-section of the population.
- Task-specific strengths: Different platforms are better/worse at different tasks: I use ChatGPT for tasks involving search and Claude for coding and iterative tasks.
- Metadata opportunities: If some anonymized user data or metadata were attached to conversations (location, device type, interaction mode (text or voice), length of tenure), we can disambiguate from the “whole economy” question: identifying industries, user sophistication, country (if it’s not too hard to negotiate privacy across jurisdictions), behavioral matters. The next step could be an anonymized unique identifier per user (similar to my project in Sierra Leone) and time stamps so you can link usage patterns. As user data is added, however, it becomes very difficult to do “blind” data analysis, like the Clio project. Even looking at descriptives may be beyond most privacy policies.
Strategic Approach: Getting a Foot in the Door
Strategically, it’s better to aim for this MVP, a mega-AEI, as a way to get a foot in the door. This could have important results for building pre-competitive collaboration.
It also parallels how the existing model of the Frontier Model Forum (FMF) operates. It’s a 501(c)(6) focused on exactly this sort of pre-competitive collaboration for AI safety, with “trusted, secure mechanisms for sharing information” modeled on cybersecurity disclosures. A parallel Economic Forum would provide a similar baseline: even while they’re competing on capabilities and market share, the labs would show willingness to work together on at least understanding the big picture. Economic effects are placed on an equal footing to safety.
One Potential Framing: Evaluation Not Metric, Dataset Not Index
Framing such an MVP as a model evaluation (eval) could be really effective. Every eval is being Goodharted: Chatbot Arena used to be a great measure, until Meta trained Llama 4 to top the chart. Evals are also contrived, separate from the questions about economic impacts we’re interested in: HLE is so far safe from Goodharting, but is made mostly of trivial questions about e.g., translating the “Biblia Hebraica Stuttgartensia”.
The framing of a Frontier Lab Economic Index as an eval could be something like: “This new model, by revealed preference of our user, is used x% more in augmenting productive work in sectors a, b, and c.” Such an evaluation would get at the economic outcomes we’re interested in much better than current capabilities-focused evaluations.
Maybe calling it an “index” is a bad idea; we don’t want a single number which can be gamed. We want a dataset, like the O*NET tasks released by Anthropic, which can be interpreted by experts into a narrative. Datapoints like “the % of chats used to augment/automate this task” are harder to Goodhart.
Finally, these data would be based on revealed preference of users in the real world, not toy problems in a sandbox like most evals.
Minimum Viable Cooperation
Are these labs willing to cooperate to the minimum level? My version of this is something like:
- A same-size sample of only chat conversations,
- From chatbots,
- Only from users on consumer plans,
- Into a black box where no other lab sees raw user data,
- For analysis by a Clio-like tool.
Complications and Considerations
Even this plan presents some complications.
What could Amazon contribute? Would we be just leaving out Meta? Both are members of the FMF. Which lab’s LLM is used for the analysis? All of them? The best open weight model? Does this MVP give Anthropic too much credit, making the other labs look like followers?
The outcome variables would have to be negotiated among the labs; the simplest would be ONET categorization across all data, or disaggregated by Lab/model. This is what was used in the AEI, as well as the best understood among economists working on automation. Sticking to ONET is great because it leaves so many potential narratives for the story. This sidesteps the problem of one model coming out as a clear “winner”.
Other ideas:
- Gimmicks (categorize e.g., existential inquiries, “how many people asked for help making soup”).
- Some measure of automation/augmentation. Anthropic had a stab at this in their second AEI release; a more generalizable method would be better.
- But we would want to generally avoid things which make one model look better/worse than others. This could be engagement metrics (avg. conversation length).
Despite the complications, something like this is a public and verifiable first step which would result in news coverage (positive, at my guess?). And importantly, it would be easy to build upon. Labs could add more models, agentic products (e.g., Copilot), time series data, and disaggregate between educational and consumer accounts.