
micro1 commits $1bn to company data for AI agents
AI training company micro1 will spend $1bn over the next 12 months buying and licensing operational data from companies. The San Francisco company announced the plan on Friday. Citi and Hercules Capital are providing the capital. The money goes through micro1’s Company Data Partnerships programme. The company uses de-identified company data to build reinforcement learning environments that reflect real business operations. In them, AI models and agents learn to work through workflows, make decisions and complete complex tasks.
micro1 said real business operations involve incomplete information, competing priorities and exceptions that require judgment. That description explains why the company is not simply buying more web text or generic labeled datasets. It wants the messy, context-heavy material that accumulates inside organizations: the standard operating procedures that tell employees what to do, the knowledge bases that capture hard-won answers, the internal documents that explain how a process actually works, the CRM records that show how customers move through a sales pipeline, the project histories that reveal what went wrong and why, and the quality assurance processes that keep output consistent.
The programme also seeks records of how teams make decisions. That could include approval chains, escalation paths, meeting notes, incident postmortems, exception logs, and human feedback on AI outputs. For AI agents, these records are valuable because they show not only what the correct answer is but also how experienced workers choose among imperfect options. An agent that can handle a refund request, an insurance claim, a supply chain disruption, or a billing dispute needs more than a rulebook. It needs to recognize when a rule should be bent, when a customer should be escalated, and when a seemingly small detail changes the entire outcome.
What micro1 pays for
The programme’s page lists the data micro1 wants. That includes standard operating procedures, knowledge bases, internal documents, CRM data, project histories and quality assurance processes. It also wants records of how teams make decisions, and human feedback on AI outputs. Payment depends on the quality, uniqueness and value of the data, the company said. The page lists three payment levels. A qualified partnership earns more than $100,000. Large datasets, or ongoing work across several teams, earn more than $500,000. Highly unique, proprietary data earns more than $1m.
Those tiers suggest micro1 is trying to price operational knowledge in a way that appeals to different kinds of enterprises. A mid-sized company with well-documented customer support workflows might qualify for the first tier. A larger business with years of project histories, structured CRM data, and multiple teams contributing ongoing documentation could reach the second tier. A company with rare, proprietary data in a specialized field, such as healthcare administration, industrial logistics, financial compliance, or legal operations, could command more than $1m. The exact amount depends on how difficult the data is to replicate and how much it improves an AI agent’s performance.
Companies keep ownership of their underlying data. micro1 said it removes personal information when relevant. It may also rewrite some datasets synthetically to loosen their link to the original records. Companies can review samples before micro1 uses the data. Those safeguards are important because operational data often contains customer names, employee details, pricing terms, legal matters, and other sensitive information. Even when data is de-identified, privacy experts warn that combining multiple datasets can sometimes re-identify individuals. Synthetic rewriting can reduce that risk, but it can also reduce the realism that makes the data valuable. micro1’s approach appears to be a balance: preserve the structure and decision logic of real work while stripping or altering identifiers and some specific details.
The programme targets companies with more than 30 employees and established documentation. micro1 gives priority to US companies, then other Western markets, and the material must be in English. Those criteria reflect both practical and legal considerations. US companies often have extensive digital records, standardized software platforms, and English-language documentation that is easy to process. Western markets may have clearer data protection frameworks, though rules vary widely. The 30-employee threshold suggests micro1 wants organizations that are large enough to have repeatable processes but not so large that negotiations become impossibly slow. Established documentation is equally important. A company with rich tacit knowledge but poor records may not be able to provide data that can be turned into reliable training environments.
Why operational data matters for AI agents
The AI industry has spent years scaling models on internet text, books, code, and other public sources. That approach produced impressive language models, but it has limits. Public data is finite, often noisy, and rarely captures the private procedures that drive real businesses. Companies increasingly want AI agents that can perform work, not just answer questions. An agent that can draft a response is useful. An agent that can resolve a customer issue, update a CRM record, follow compliance rules, and know when to ask a human for help is far more valuable. Building that kind of agent requires training and evaluation environments that resemble the real world.
Reinforcement learning is one method for teaching agents through trial and error. In a simulated environment, an agent takes actions, receives feedback, and adjusts its behavior. If the environment is unrealistic, the agent learns the wrong lessons. It may exploit shortcuts, ignore exceptions, or fail when information is missing. micro1’s environments are designed to reflect real business operations, including incomplete information, competing priorities and exceptions that require judgment. That makes the training harder, but it also makes the resulting agents more robust. A customer service agent trained on real refund policies, escalation rules, and customer histories will be better prepared for the ambiguous cases that cause most problems.
Operational data can support many domains. In finance, agents could learn to reconcile transactions, flag suspicious activity, and prepare compliance reports. In healthcare, they could help with prior authorization, scheduling, and claims processing, provided privacy rules are strictly followed. In manufacturing and logistics, they could respond to supply chain disruptions, adjust maintenance schedules, and coordinate inventory. In software, they could triage bugs, update documentation, and automate incident response. In human resources, they could answer policy questions, manage onboarding, and route employee concerns. Each domain has its own vocabulary, rules, exceptions, and risk tolerances. Generic models often struggle with those specifics. Domain-specific operational data gives agents the context they need.
The financing and the market for data
Citi and Hercules Capital are providing the capital for the $1bn commitment. Citi is a global bank with deep ties to corporate finance. Hercules Capital is a specialty lender that works with venture-backed and growth-stage companies. Their involvement signals that data acquisition is being treated as a financeable business expense, not just a research project. If micro1 can convert licensed company data into better AI training environments, it may be able to sell those environments or the resulting models to enterprises and AI labs. The $1bn figure is large enough to move markets. It suggests that demand for high-quality, private operational data is strong enough to justify a multi-year investment.
The market for AI training data has become increasingly competitive. Public web data is scraped at scale, but copyright lawsuits and changing platform policies have made it riskier to rely on. Some AI companies have signed licensing deals with publishers, stock image providers, and social platforms. Others have turned to synthetic data, human feedback, and expert annotations. Operational data from companies is another frontier. It is not available on the open web. It often contains the kind of tacit knowledge that gives a business its competitive edge. That makes it scarce, defensible, and potentially very valuable.
micro1 calls itself a data lab that helps AI research companies and enterprises train frontier models and evaluate agents. It runs three in-house labs, one of them for robotics. That background matters because collecting data is only the first step. The harder work is structuring it, cleaning it, simulating workflows, designing reward signals, and evaluating whether an agent behaves safely and accurately. A company that only brokers data may struggle to deliver useful training environments. micro1’s positioning suggests it wants to control more of that pipeline, from data partnership to environment design to agent evaluation.
Competition for proprietary records
In September, micro1 made a competing bid for Spirit Airlines’ customer records. The court-appointed privacy official in Spirit’s bankruptcy has since backed Google’s $10m deal for that data. The episode shows how valuable customer and operational records can become when a company fails or restructures. Bankruptcy estates may hold datasets that are difficult to recreate, including years of customer interactions, loyalty behavior, and service histories. Privacy officials and courts must weigh the value of those records against the privacy rights of the people in them. Google’s winning bid also shows that major technology companies are willing to pay for data that can improve their AI systems.
Elon Musk’s SpaceX has also weighed buying data from startups to train its AI models. That report, along with micro1’s program, points to a broader trend: AI developers are looking beyond public web scraping and toward private, proprietary, and operational data. Startups may sell data to raise capital. Large enterprises may license data to create new revenue streams. Data brokers may emerge to connect the two sides. Regulators may struggle to keep up, especially when data crosses borders or is transformed synthetically.
For companies considering micro1’s programme, the decision is not only about money. It is about control, reputation, and competitive risk. Licensing operational data could fund new initiatives and improve internal AI projects. It could also help a competitor build a better agent. Even with ownership retained and personal information removed, companies must think about what their data reveals about their processes, pricing, weaknesses, and institutional knowledge. The payment tiers may be attractive, but the strategic implications may be just as important. micro1’s $1bn commitment will test whether businesses are ready to treat their operational knowledge as a product.
