S&P 500 7,743.41 +0.51%Nasdaq 27,068.72 +0.48%Dow 51,828.62 +0.93%Russell 2000 2,837.55 +0.07%as of 2026-09-25 close
◈ Frontier Tech Wire
Quantum, AI and frontier-tech small caps — on the wire
Funding & Deals

Gimlet Labs Raises $300m at a $3bn Valuation With Arm and Microsoft's M12. The 10x Speed Figure Is Scoped to One Configuration

Reports on Friday put a $300 million round at a roughly $3 billion post-money valuation, led by Andreessen Horowitz. The speed multiples attached to the news trace back to two March documents that state them with very different levels of qualification.
Illustrative photograph: people working in a business setting.

Gimlet Labs, the inference-orchestration startup founded by the team behind the Kubernetes observability company Pixie, has raised about $300 million at a post-money valuation of roughly $3 billion, according to reporting published Friday by Bloomberg and picked up across the startup-funding trade press. Andreessen Horowitz led the round. Arm Holdings and M12, Microsoft's venture arm, are described as new investors.

Two details are worth establishing before anything else. First, as of Friday morning the company had not published its own announcement: the most recent funding post on the Gimlet Labs blog is still the March 23 Series A note, and the blog's newest entries are engineering write-ups dated June 29 and July 8. A newswire search returns a Gimlet release for the March Series A and none for this round. The round is being reported rather than released. Second, the analysis outlet FourWeekMBA, summarising the same reporting, states plainly that revenue, customer count and traction were not disclosed. Anyone modelling a $3 billion mark is doing so without those inputs.

What Gimlet sells is easier to describe than to price. Menlo Ventures, which led the company's Series A, characterised it in a March 23 investment note as a multi-silicon inference and compute cloud: software that decomposes an inference request into stages and routes each stage to whichever silicon suits it, mixing GPUs, SRAM-centric accelerators and CPUs without asking the developer to rewrite anything. The pitch is vendor-neutrality as a product, which is why the identity of the two strategic backers reads as a thesis statement. Arm's commercial interest and Microsoft's M12 both sit on the side of a world in which inference is not welded to one architecture.

The funding history is compressed. Per TechFundingNews, the company raised a $12 million seed from Factory alongside angels including Intel chief executive Lip-Bu Tan, Figma's Dylan Field and a16z general partner Raghu Raghuram, then an $80 million Series A led by Menlo Ventures with Eclipse, Factory, Prosperity7 and Triatomic. That Series A was announced on March 23. Friday's round therefore lands roughly five and a half months later at a valuation the earlier round did not imply.

The number doing the most work in the coverage is a performance claim: TechFundingNews reports three-to-ten-times inference speed improvements for the same cost and power envelope. That range, in that form, appears in neither of the two March documents this desk was able to open. What those documents do say is narrower, and the difference between them is the point.

The first is a Gimlet engineering post dated March 11, titled "Low-Latency Inference with Speculative Decoding on d-Matrix Corsair and GPU." That post reports, in its own words, "2-5X end-to-end request speedup on configurations optimized for interactivity, and up to 10X end-to-end speedup for energy-optimized configurations." It frames the same result a second way, as a "2-10X reduction in end-to-end request latency versus running the same speculative decoder on GPU."

Those figures come with a specification. The target model is gpt-oss-120b with a 1.6-billion-parameter draft model, on an 8,000-token input and 1,000-token output, framed as a coding workflow. The accelerator is d-Matrix's Corsair part, which the post describes as carrying 2GB of on-chip SRAM and 150 TB/s of memory bandwidth. The comparison point is a latest-generation GPU whose exact SKU the post deliberately omits, on the stated grounds that the point of the comparison is the architectural trade-off rather than the part. The post also notes that the advantage widens as more tokens are verified per step, because fast drafting on Corsair lowers the cost of rejected tokens.

The second document compresses that work. d-Matrix's announcement of the partnership, dated March 12, is headlined "d-Matrix and Gimlet Labs to Deliver 10x Speed Ups, Massive Power Efficiency for Frontier AI Workloads," and claims "10x performance benefits in latency and throughput per Watt compared to GPU-only approach," describing order-of-magnitude increases on both inference latency and throughput per watt. The quantity is not where the two documents part company — both put a 10x on wall-clock latency. The scope is. Gimlet's own post reserves the up-to-10x figure for energy-optimised configurations and puts interactivity-optimised configurations at 2x to 5x, on one named model, one sequence length and one accelerator pairing against an unnamed GPU. The vendor release carries the 10x without those conditions attached and points to no third-party validation. Both numbers are sourceable; only one of them is qualified, and it is the qualified one that came first.

Gimlet's other published work is narrower in scope and specific about hardware. The blog index lists a benchmarking post on AI-generated CUDA kernels on an H100 dated October 18, 2025, one on AI-generated Metal kernels for Apple devices dated August 26, 2025, and one on splitting LLM inference across different hardware platforms dated October 13, 2025. The multiples reported inside those posts were not re-verified for this piece and are not repeated here.

The company has also been extending beyond software. Its June 29 post describes work with MLCommons on benchmarks for agentic workloads, and a July 8 post covers formal verification of AI-generated GPU kernels — a reasonable place to invest if your business depends on customers trusting machine-written kernels in production. Coverage of Friday's round says Gimlet has moved into data-centre configuration and infrastructure deployment as well, which would put it in a different cost structure than a pure software layer.

The founding team is the Pixie team: Zain Asgar as chief executive, with Michelle Nguyen, Omid Azizi, Natalie Serrino and James Bartlett. Pixie was acquired by New Relic in 2020. In the March 12 d-Matrix release, Asgar framed the pairing as sending to d-Matrix's hardware the phases of inference on which GPUs waste energy, and d-Matrix chief executive Sid Sheth framed it in terms of power limits capping how fast AI can advance.

For a desk that covers the money moving around AI and specialised silicon, the round is a marker of where capital is going: not into another accelerator, but into the layer that decides which accelerator runs what. That is a bet that the inference market stays heterogeneous. It is also, at $3 billion, a bet placed without published revenue.

What would move this story from reported to documented is straightforward: a company release naming the round, the security sold and the participating investors; any disclosure of customers or run-rate; and a benchmark run on named hardware by someone other than the two vendors whose products are being compared. None of those had surfaced in the documents this desk was able to open as of Friday morning.

This article is for general information only and is not investment advice. Figures are as reported by the cited sources at time of writing.

Related coverage