Smaller AI models put the spend back where partners earn margin
Marta Garcia told Raconteur that AI costs keep climbing as organizations expand their use of it. On-device inference moves the money to hardware and services UK partners can sell

- Multiverse Computing’s CFO says customers buy compression to cut infrastructure cost and to run models locally on laptops, phones and vehicles instead of the cloud.
- The company’s own releases claim 80–95% smaller models, but its 24 September speech-to-text test also shows the word error rate rising when a model is halved.
- Edge inference turns AI from a consumption credit the hyperscaler prices into a hardware refresh, an endpoint management job and a deployment project partners can invoice.
Marta Garcia, chief financial officer of Multiverse Computing, a San Sebastián company that shrinks AI models, put the problem plainly in an interview with Raconteur published on 24 September. “As organizations expand their use of AI, costs continue to increase,” she told Raconteur. That is the whole AI economy in 11 words, and it is the sentence UK channel leaders should be planning around for 2027.
Garcia joined the company when it had seven or eight staff and told Raconteur it now has around 500. Her account of what customers actually buy is more useful than most vendor pitches. Governments, pharmaceutical companies and financial institutions, she said, come to Multiverse to cut infrastructure cost. Device makers want models that run efficiently on the hardware they ship. In Europe, data sovereignty sits alongside both, particularly around AI gigafactories. The company’s compressed models use less energy, she said, and can be squeezed far enough to run on a laptop, a phone, a vehicle or a drone, so the cloud carries only part of the load.
She named two markets the company is chasing with its current fundraising: gigafactories, where European investment is heavy, and edge deployment. She also described a new product built in-house: a small language model that lives on the device and works as a router, triaging each request so the easy ones stay local and the rest go to the cloud.
Strip out the quantum branding and the thesis is this: once AI spend is measured in inference rather than experiments, the bill scales with usage and efficiency becomes the constraint. She is right, and the consequence for UK partners is bigger than the company itself.
The vendor’s numbers are large and its own
Multiverse’s figures deserve the usual caution. In its Series C announcement on 27 July, the company said it was targeting up to $570m (£430m) at a pre-money valuation of $1.7bn, five times the Series B valuation, and that it now has more than 100 customers. The same release says CompactifAI cuts the size of large language models by up to 80–95% with what it calls “immaterial accuracy loss”. Its Series B release in June 2025 claimed a 50–80% reduction in inference costs. These are the vendor’s numbers, not independent tests. CB Insights put Multiverse in its AI 100 for 2025, in the infrastructure category, which tells you analysts rate the technology and says nothing about any particular customer’s bill.
The more telling release came on 24 September, the same day as the interview. Multiverse said, in a release issued with HPE and Intel, that a compressed version of the Whisper speech-to-text model halved its parameters, from 0.8 billion to 0.4 billion, and transcribed about 17,000 hours of audio a day on a single HPE ProLiant server with two Intel Xeon 6 processors and no graphics processing unit (GPU). Throughput roughly doubled. The awkward line is further down: the word error rate rose from 2.14% to 2.75%. Compression has a price. For a contact center it is a rounding error; for a regulated workflow it is a conversation with the compliance team.
Shrinking models move the spend to the endpoint
Follow the spend. Cloud inference is billed by the token through a hyperscaler, and the UK partner’s share of that is a resale margin on consumption credits the hyperscaler sets and can cut. On-device inference is different. It is a hardware refresh, because a model that fits in 0.75GB of memory still needs a machine built to run it. It is an endpoint management problem, because a fleet of laptops each running a local model needs patching, version control and a policy for what the router sends upstream. It is a deployment project, because someone has to decide which use cases stay local, test the accuracy loss and sign it off. Every one of those is a line partners already invoice.
That is the shift to plan around. According to CB Insights, funding to AI companies passed $170bn between the start of 2024 and April 2025. CB Insights says the bulk of it went to the largest model builders, which means the data centers that serve them. The next phase is paying to run them, and the customers Garcia describes, the finance director at a bank or a council, are the ones who notice that bill first. Their AI budgets for 2027 will be set on 2026 cloud invoices, and those invoices are why efficiency vendors are raising money at unicorn valuations.
The router model is the tell
Garcia’s router model is the detail to watch. A device that decides for itself what to keep local and what to send upstream means the cloud does not go away; it becomes the overflow. For partners that is the best of both. The consumption resale stays, at a smaller scale, and around it sits the hardware, the management layer and the services that decide the split. Partners who sell only the credits will find the credits shrinking. Partners who sell the split will be paid for judgment.
Three things follow. First, get AI-capable laptops and edge servers into the 2027 refresh conversation now, with a compressed-model demonstration rather than a vendor slide, because the customer will want to see what the accuracy loss looks like on their own data. Second, treat model version control as part of endpoint management, and price it. Third, be honest about the trade-off. Multiverse’s own test shows an error rate climbing when a model is halved, and a partner who says so before the customer finds out will keep the account.
Garcia’s line about rising costs is a warning to her own customers. For the UK channel it is closer to an invitation. The hyperscalers made AI a bill you rent. Efficiency is making it something you install, manage and support again, and that has always been where partners earn their keep.
Get The VETTDD BriefingThe week in the technology channel, every week.
Subscribe free- Raconteur, “CFO on the Spot: Five minutes with Marta Garcia, CFO of Multiverse Computing”, by Rayanne Harmon, 24 September 2026. https://www.raconteur.net/finance/cfo-on-the-spot-five-minutes-with-shilpa-kaluti-cfo-of-scrumconnect
- Multiverse Computing, “Multiverse Computing Announces Series C Fundraising Targeting up to $570M (€500M) to Power Efficient AI from Edge to Cloud”, press release, 27 July 2026. https://multiversecomputing.com/resources/multiverse-computing-announces-series-c-fundraising-targeting-up-to-usd570m-eur500m-to-power
- Multiverse Computing, “Multiverse Computing, HPE and Intel run enterprise speech-to-text on CPUs alone”, press release, 24 September 2026. https://multiversecomputing.com/resources/multiverse-computing-hpe-and-intel-run-enterprise-speech-to-text-on-cpus-alone
- Multiverse Computing, “Multiverse Computing Raises $215M to Scale Ground-Breaking Technology that Compresses LLMs by up to 95%”, press release, 12 June 2025. https://multiversecomputing.com/resources/multiverse-computing-raises-usd215m-to-scale-ground-breaking-technology-that-compresses-llms-by
- CB Insights, “AI 100: The most promising artificial intelligence startups of 2025”, report, 24 April 2025. https://www.cbinsights.com/research/report/artificial-intelligence-top-startups-2025/
- Multiverse Computing, team page (Marta García, chief financial officer). https://multiversecomputing.com/team



