Our friends at CoreWeave have been busy. Over the last few weeks, they published a two-part series on CoreWeave AI Object Storage (Part 1, Part 2), plus benchmark posts showing 2+ GB/s per GPU and then 7+ GB/s per GPU. The star of all of them is LOTA, the Local Object Transport Accelerator: a caching proxy that runs on every node in a CoreWeave Kubernetes Service (CKS) cluster and keeps objects on the node's local NVMe.
I want to be upfront about something. I like that they published their methodology. They used Warp, which is our own open-source S3 benchmark tool, and they told us the object sizes, the concurrency, and the node configuration. That's exactly how this industry should talk about performance, and it's what makes an engineering conversation like this one possible.
So let's have that conversation. My read of their numbers is simple: a cache exists because the thing behind it is slow. When your object store is fast, you don't need to make complicated trade-offs to keep GPUs fed.
What the numbers actually say
Let's start with CoreWeave's 2+ GB/s per GPU benchmark, since it's the one with the most detail. They ran Warp on 20 GPU nodes, each with 8 GPUs, 1 TiB of LOTA cache per node, and a dedicated network adapter with dual 100 Gbps links for storage access. Each node wrote 10,000 objects of 50 MiB, then read them back randomly for 10 minutes.
Here's what they reported:
- While the cache was filling (cold): about 24 GiB/s across the whole cluster, or 1.2 GiB/s per node.
- Once the cache was warm: 368 GiB/s across the cluster, or 18.4 GiB/s per node, which they divide by 8 GPUs to get 2.3 GiB/s per GPU.
Their own write-up says it plainly: the performance of cached or pre-staged data "is entirely driven by the performance of the GPU-node-local storage." In other words, the headline number is a measurement of the local NVMe drives inside the GPU servers. The object store shows up in the cold number, and the cold number is 1.2 GiB/s per node.
The 7+ GB/s per GPU result follows the same pattern. It was measured on 16 nodes of a newer GPU server generation, with a higher-speed interconnect and pipelining, at about 28 GB/s per node, then divided by 4 GPUs per node instead of 8. CoreWeave is transparent about this: they attribute the first doubling (2 to 4 GB/s per GPU) to having fewer GPUs per node, the next jump to moving cache-to-cache traffic onto that faster interconnect, and the rest to LOTA pipeline improvements. Two more details worth knowing: these are warm-cache GET tests, and the 2+ GB/s per GPU post says the I/O was CPU-driven.
None of that is wrong. It's just measuring a cache, not an object store.
So how fast can an object store be without a cache?
Here's a benchmark from The New Economics of AI Factory Efficiency, a public joint whitepaper from Solidigm, MinIO and Intel. It ran AIStor with the same tool (Warp) and no caching layer at all. Every GET is served by the durable, erasure-coded object store over plain TCP.
The cluster:
- 8 storage nodes, each with a single Intel Xeon 6781P, 504 GB of DDR5 and two 400GbE NICs.
- 24 Solidigm D5-P5336 122.88 TB drives per node. These are QLC drives, built for capacity. That's about 2.95 PB raw per node and roughly 24 PB for the cluster.
- Erasure coding EC:4 across 24 erasure sets of 8 drives each, so every object is written as four data shards and four parity shards.
- 8 client nodes running Warp, each sending test traffic over one 400GbE NIC.
- 256 MiB objects, 32 concurrent connections per client, and standard TCP transport.
The results (16 phases, 8 PUT and 8 GET, stepping from 1 to 8 clients, 15 minutes each, 240 minutes of continuous I/O, zero errors):
- GET: 267.8 GiB/s across the cluster (peak median at 8 clients). The fastest second hit 272.4 GiB/s, and the cluster still wasn't read-bound.
- PUT: 120 GiB/s across the cluster, reached at 7 clients and flat at 8.
- Median first-byte latency stayed at 3 to 5 ms across the whole sweep, and even the worst second of the 8-client GET run was within about 11% of the median.
That works out to roughly 33.5 GiB/s of durable GET per storage node, from QLC flash, with parity, over plain TCP. Put that next to CoreWeave's numbers:

I want to be careful here, because this is exactly where benchmark posts go wrong. These aren't identical tests. CoreWeave's per-node numbers are per GPU node reading from its cache; ours are per storage node serving clients. The object sizes and client counts are different too. You can't draw a single bar chart and declare a winner.
But you don't need one to see the point. An erasure-coded object store on capacity-optimized QLC drives, running over plain TCP, delivers more throughput per node than a warm cache on local NVMe (about 33.5 GiB/s versus 18.4 GiB/s), and more than 11 times the cluster throughput of CoreWeave's own cold path. A cache should be faster than what's behind it. When a durable store with parity beats the cache, the architecture is telling you something.
And this cluster isn't even built for speed. It uses QLC drives, it's configured for durability with as many parity shards as data shards, and it runs on standard TCP networking.
The cost of the cache
Part 2 of CoreWeave's series, along with their LOTA and Warp docs, is a practical guide to getting good performance out of LOTA, and it's worth reading for what it asks you to manage:
- Point your clients at a different endpoint (cwlota.com instead of cwobject.com) when running inside the cluster.
- Pre-stage your dataset by issuing a zero-length-range HeadObject call against every object before the job starts, so the first epoch isn't a string of cold misses.
- Keep objects above 4 MB, because anything smaller isn't cached and goes straight to persistent storage. Use multipart uploads for large objects.
- When benchmarking, start with 15 MiB objects "to stay above the threshold where metadata overhead becomes significant."
- Plan on each node contributing 1 TiB of local NVMe to the cache by default.
- LOTA accelerates reads only. Writes are a separate problem, handled by a separate cross-region write feature.
Every one of those is a reasonable thing to do. Together they're a list of trade-offs you only have to make because the backing store can't feed the GPUs on its own. It's also CKS-only: the cache lives inside CoreWeave's Kubernetes service, on their nodes.
To be clear, I'm not arguing against caching data close to GPUs in other regions. If your datasets live in one data center and your GPUs in another, you want something like a CDN for objects, and CoreWeave has thought hard about that. That's a data placement choice, and a fair one. My argument is narrower: you shouldn't need a cache on every GPU node to make your object store fast in the first place.
When the data is KV cache, use a KV cache
CoreWeave's 7+ GB/s per GPU post says CAIOS "powers high-throughput key-value caches for inference." That's worth a quick word, because it mixes up two different layers.
An object cache speeds up S3 GETs of datasets and checkpoints. An inference KV cache stores attention state so the engine doesn't recompute a long prompt every time a user comes back. They have different access patterns, block sizes, and success metrics. Warp GET throughput doesn't tell you anything about time to first token.
That's why we built MemKV as its own tier: a shared-nothing KV store on raw NVMe, wired into inference engines like vLLM. In our benchmark serving a 500B+ parameter model in FP8 on a single 8-GPU node, with an agentic coding workload and 30 to 60 second think times between turns:
- P99 time to first token for returning users improved up to 4.2× (143.6 s to 34.0 s at concurrency 128).
- Throughput per GPU was 45% higher than the baseline at that same concurrency.
- Zero transport errors and zero missing keys across 6.4 TiB written.
We measure that with an inference benchmark, because that's the workload. Right tool, right layer, right benchmark.
What we'd like to see
We're not asking anyone to take our word for it. Here's what would move this conversation forward:
- Publish cold and warm side by side. Report CAIOS primary-endpoint throughput (cwobject.com, no cache) next to LOTA-warm numbers, with the same Warp settings. The cold number is the object store.
- Show per node, not just per GPU. Dividing by GPU count makes the number move with the server model rather than the storage.
- Publish writes. The 2+ GB/s per GPU post said write performance was next. Checkpoints are where GPUs sit idle.
- Benchmark each layer with its own workload. Warp for object storage, an inference benchmark that measures time to first token for KV cache.
We'll hold ourselves to the same standard. Warp is open source; the full hardware configuration and tuning for the whitepaper run are published, and we're happy to share configs with anyone who wants to reproduce it.
If you're thinking about how to keep your GPUs fed, reach out to us. We'd love to show you how fast an object store can be when it doesn't need a cache to keep up.




