Performance

MinIO Blog Posts

You Don’t Need a Cache. You Need a Faster Object Store
arrow
CoreWeave's LOTA benchmarks are transparent and well run, but the warm-cache headline numbers measure local NVMe inside the GPU nodes, not the object store behind them. Using the same open-source Warp tool with no cache at all, MinIO AIStor delivered roughly 33.5 GiB/s of durable, erasure-coded GET throughput per storage node over plain TCP, making the case for a faster store rather than another layer in front of it.
AIStor
AI/ML
Performance
Storage & Infrastructure
Prompt Caching: Stop Paying GPUs to Read the Same Prompt Twice
arrow
Every coding assistant, enterprise chatbot, and agent loop resends the same tool schemas, system instructions, and policy text on every turn, and the GPU rebuilds all of it into attention state before it can emit a single output token. Prompt caching computes that prefix once and reuses the KV state, and MemKV takes it from a per-process optimization to a shared NVMe-backed tier that survives routing across replicas, HBM eviction, worker restarts, and concurrency.
Agent Memory
AI/ML
AIStor
Performance
Operations
From Cache Hits to Production SLAs | Part 3 of 3
arrow
Cache hits are not the outcome. Part 3 turns the architecture from Parts 1 and 2 into an evaluation framework: capture an honest tail-latency baseline first, treat MinIO's published 53-second-to-703-millisecond TTFT result as a proof point to reproduce rather than a business-case input, and follow the measurement chain from repeated context through prefix reuse and avoided recompute to unit economics, ending in a buyer checklist that judges a shared context tier on P99 TTFT and jitter rather than throughput or capacity.
Agent Memory
AI/ML
AIStor
Operations
Performance
When Repeated Context Becomes an Infrastructure Problem | Part 2 of 3
arrow
A prefix cache that only helps one process is useful, but requests move across replicas, HBM fills, sessions spill, and workers restart, so reuse that lives inside a single worker is not a fleet architecture. Part 2 works through what the serving stack needs once KV state leaves local GPU memory: a tier that is larger than HBM, fast enough that restore beats recompute, shared across workers, and reachable through the runtime's own KV transfer path, which is memory behavior at cluster scope rather than storage.
Agent Memory
AI/ML
AIStor
Performance
Prompt Caching Is an AI Margin Lever, Not a Model Trick | Part 1 of 3
arrow
Agentic AI applications resend the same project rules, tool schemas, and document context on every turn, so an expensive GPU fleet spends much of its time rebuilding a prefix it has already processed. Part 1 of three reframes prompt caching as an operating-margin lever rather than a model feature, maps prompt caching, prefix caching, KV cache, and KV cache offload to the business questions each one answers, and argues that reusable context needs a memory path rather than ordinary enterprise storage.
Agent Memory
AI/ML
AIStor
Performance
MinIO text on dark background: The Complete GPU Storage Stack: AIStor + MemKV Architecture Guide.
The Complete GPU Storage Stack: AIStor + MemKV Architecture Guide
arrow
Learn how AIStor, MemKV, and a secure data fabric form the three-layer GPU stack that slashes inference latency.
Architecture & Design Patterns
MemKV
Performance
Glowing white cube among rows of translucent blue cubes in a digital 3D pattern.
Search Compressed Data Without Decompressing It
arrow
Learn how MinLZ enables fast, selective searches of compressed data without decompression, dramatically reducing I/O and accelerating queries on object storage.
Performance
Logos of MinIO, Solidigm, and Intel on a dark background with pink and orange cloud-like shapes.
Density Was Supposed to Cost You Performance. The Numbers Say Otherwise.
arrow
Solidigm and MinIO put high-density QLC NVMe object storage under real load and disproved the long-held assumption that density costs performance. The tested node is the same building block that scales directly into MinIO's ExaPOD reference architecture, so the path from a single pod to exascale runs on identical hardware and software.
AI/ML
Performance
Storage & Infrastructure
Architecture & Design Patterns
Swirling abstract smoke in purple, pink, and orange hues on a dark background.
GPU-Accelerated Semantic Search with NVIDIA cuVS and MinIO AIStor
arrow
Not every semantic search problem needs a vector database. NVIDIA cuVS runs GPU vector search in memory; MinIO AIStor persists every artifact from raw documents to indexes. Full pipeline, working code.
AI/ML
Performance
Introducing MinIO MemKV: Purpose built Context Store for Inference at scale
arrow
MinIO MemKV eliminates the recompute tax in GPU inference clusters with shared petabyte-scale context memory.
AI/ML
Cloud Infrastructure
Performance
Storage & Infrastructure
MemKV
Benchmarking Vector Index Creation with MinIO AIStor, Milvus, and NVIDIA cuVS
arrow
106 million vectors indexed 12x faster. AIStor with NVIDIA cuVS and GPUDirect RDMA rewrites the benchmark.
AI/ML
AIStor
Performance
Integrations & Partners
Storage & Infrastructure
Long corridor in a data center with rows of illuminated servers and blue lighting reflections.
AIStor Inside NVIDIA BlueField-4: Object Data at Wire Speed
arrow
Legacy storage talks to the AI factory. AIStor lives inside it — running natively on NVIDIA BlueField-4 Vera
AIStor
AI/ML
Performance
Storage & Infrastructure
Architecture & Design Patterns
Hadoop logo with a small yellow elephant icon and blue stylized text on a white background.
Hadoop HDFS's Logical Successor
arrow
MinIO is Hadoop HDFS's logical successor—faster, cloud-native, S3-compatible with better economics & simplicity
Performance
Data Lakes & Analytics
Case Studies & Solutions
Apache Ecosystem
Close-up of rows of smooth, spherical balls with shallow depth of field in black and white.
The Small Files Problem: Solutions for Big Data
arrow
Inline metadata and erasure coding solve small file performance challenges at scale
Architecture & Design Patterns
Data Lakes & Analytics
Performance
Text reading Unlocking AI/ML Performance with MinIO and AMD logos on a white background.
Unlocking AI/ML Performance with AMD + MinIO
arrow
MinIO + AMD: Unlocking AI/ML performance with EPYC processors & NVMe drives for high-throughput workloads
AI/ML
Integrations & Partners
AIStor
Performance
Close-up of colorful glass marbles scattered on a dark surface with soft lighting and shadows.
Working with Small Objects in AI/ML workloads
arrow
How MinIO handles small file performance with erasure coding and inline metadata
Performance
Operations
Abstract black and white photo showing diagonal light and shadow with striped pattern on a wall.
Databases on Object Storage - the New Normal
arrow
Databases on object storage—Druid, ClickHouse, Snowflake, Teradata leverage MinIO for disaggregated scale & performance
Performance
Architecture & Design Patterns
Data Lakes & Analytics
Car dashboard with illuminated speedometer, tachometer, fuel gauge, and warning lights at night.
Benchmarking AIStor with WARP and Perf test
arrow
Guide to benchmarking MinIO performance with WARP tool measuring throughput and latency
Operations
Performance
Colorful digital lines flowing into a block of glowing binary code on a dark grid background.
Data and Drive parity on AIStor
arrow
Flexible erasure code configuration for custom data and parity drive ratios per deployment
Performance
Operations
AIStor
Highway at night with long exposure showing white and red light trails of moving vehicles.
WARP speed your AI data storage Infrastructure
arrow
WARP-speed AI data storage—measure MinIO performance with distributed S3 benchmarking for ML infrastructure validation
AI/ML
Performance
Architecture & Design Patterns
Storage & Infrastructure
Digital Earth at night with networked glowing nodes and blue lines forming a three-dimensional grid.
MinIO Networking with Overlay Networks
arrow
Overlay networks with MinIO—Docker & Kubernetes CNI enable secure multi-host container communication for distributed apps
Operations
Kubernetes & Containers
Performance
Abstract digital servers connected by glowing red data streams in a futuristic grid environment.
Replication Strategies Deep Dive
arrow
Replication strategies deep dive—compare batch replication, mc mirror, site & bucket replication for different use cases
Operations
Architecture & Design Patterns
Kubernetes & Containers
Cloud Infrastructure
Performance
3D abstract cube made of smaller interconnected glowing cubes linked by light blue beams.
Scaling up MinIO Internal Connectivity
arrow
Internal cluster network architecture and best practices for distributed MinIO scaling
Operations
Performance
Blue connected 3D cubes linked by white parallel lines on a gradient blue background.
Day 2 with MinIO: Scaling, Hardware Ops, Administration
arrow
MinIO Day 2 operations—administration, monitoring, scaling, updates & troubleshooting for production cluster management
Operations
Architecture & Design Patterns
Performance