AI/ML

MinIO Blog Posts

You Don’t Need a Cache. You Need a Faster Object Store
arrow
CoreWeave's LOTA benchmarks are transparent and well run, but the warm-cache headline numbers measure local NVMe inside the GPU nodes, not the object store behind them. Using the same open-source Warp tool with no cache at all, MinIO AIStor delivered roughly 33.5 GiB/s of durable, erasure-coded GET throughput per storage node over plain TCP, making the case for a faster store rather than another layer in front of it.
AIStor
AI/ML
Performance
Storage & Infrastructure
MinIO Doubles the Sessions Your GPUs Hold, Measured the Way You'd Actually Run It
arrow
Published KV cache speedup claims run from sixteen to seventy times, but they come from tests that disable the GPU's own cache and drive traffic from a load generator with no human pauses, which measures a component rather than a fleet. MinIO rebuilt the benchmark around real agentic session behavior, and on the same GPUs MemKV delivered 2.4× the tokens, 2.59× the completed turns, and 4.5× better tail latency on returning turns.
Agent Memory
AI/ML
Storage & Infrastructure
Prompt Caching: Stop Paying GPUs to Read the Same Prompt Twice
arrow
Every coding assistant, enterprise chatbot, and agent loop resends the same tool schemas, system instructions, and policy text on every turn, and the GPU rebuilds all of it into attention state before it can emit a single output token. Prompt caching computes that prefix once and reuses the KV state, and MemKV takes it from a per-process optimization to a shared NVMe-backed tier that survives routing across replicas, HBM eviction, worker restarts, and concurrency.
Agent Memory
AI/ML
AIStor
Performance
Operations
From Cache Hits to Production SLAs | Part 3 of 3
arrow
Cache hits are not the outcome. Part 3 turns the architecture from Parts 1 and 2 into an evaluation framework: capture an honest tail-latency baseline first, treat MinIO's published 53-second-to-703-millisecond TTFT result as a proof point to reproduce rather than a business-case input, and follow the measurement chain from repeated context through prefix reuse and avoided recompute to unit economics, ending in a buyer checklist that judges a shared context tier on P99 TTFT and jitter rather than throughput or capacity.
Agent Memory
AI/ML
AIStor
Operations
Performance
When Repeated Context Becomes an Infrastructure Problem | Part 2 of 3
arrow
A prefix cache that only helps one process is useful, but requests move across replicas, HBM fills, sessions spill, and workers restart, so reuse that lives inside a single worker is not a fleet architecture. Part 2 works through what the serving stack needs once KV state leaves local GPU memory: a tier that is larger than HBM, fast enough that restore beats recompute, shared across workers, and reachable through the runtime's own KV transfer path, which is memory behavior at cluster scope rather than storage.
Agent Memory
AI/ML
AIStor
Performance
Prompt Caching Is an AI Margin Lever, Not a Model Trick | Part 1 of 3
arrow
Agentic AI applications resend the same project rules, tool schemas, and document context on every turn, so an expensive GPU fleet spends much of its time rebuilding a prefix it has already processed. Part 1 of three reframes prompt caching as an operating-margin lever rather than a model feature, maps prompt caching, prefix caching, KV cache, and KV cache offload to the business questions each one answers, and argues that reusable context needs a memory path rather than ordinary enterprise storage.
Agent Memory
AI/ML
AIStor
Performance
We deleted the agent mid-sentence. The work continued.
arrow
Worker A gathers evidence, publishes an accepted handoff, starts another edit, and is deleted mid-sentence with the unfinished tail left visible. Worker B starts in a fresh runtime with no session state, verifies the last accepted boundary in the same authorized AIStor Memory Workspace, discards the unchecked tail, and continues the work rather than restarting it.
Agent Memory
AI/ML
AIStor
Your inbox agent has no business remembering your workouts
arrow
Personal AI agents become useful as they learn you, but that familiarity should not require one agent accumulating your entire life. AIStor Memory gives each agent a bounded relationship with its own learned history, active workspace, and credential scope, all under your control.
AI/ML
Agent Memory
AIStor
Introducing AIStor Memory: Long-Term Memory For AI Agents
arrow
Every agent begins with the experience your organization has already earned.
AI/ML
AIStor
Integrations & Partners
Agent Memory
Logos of MinIO, Solidigm, and Intel on a dark background with pink and orange cloud-like shapes.
Density Was Supposed to Cost You Performance. The Numbers Say Otherwise.
arrow
Solidigm and MinIO put high-density QLC NVMe object storage under real load and disproved the long-held assumption that density costs performance. The tested node is the same building block that scales directly into MinIO's ExaPOD reference architecture, so the path from a single pod to exascale runs on identical hardware and software.
AI/ML
Performance
Storage & Infrastructure
Architecture & Design Patterns
Swirling abstract smoke in purple, pink, and orange hues on a dark background.
GPU-Accelerated Semantic Search with NVIDIA cuVS and MinIO AIStor
arrow
Not every semantic search problem needs a vector database. NVIDIA cuVS runs GPU vector search in memory; MinIO AIStor persists every artifact from raw documents to indexes. Full pipeline, working code.
AI/ML
Performance
Introducing MinIO MemKV: Purpose built Context Store for Inference at scale
arrow
MinIO MemKV eliminates the recompute tax in GPU inference clusters with shared petabyte-scale context memory.
AI/ML
Cloud Infrastructure
Performance
Storage & Infrastructure
MemKV
What Happens When Databricks Can Query Your On-Premises Data Directly
arrow
Until recently, data that stayed on-premises was data that Databricks couldn't reach. If your analytics and AI workloads ran in Databricks, and your most valuable data lived on-prem, you had two options: build and maintain a replication pipeline to copy data into the cloud, or accept that certain datasets simply wouldn't participate in your cloud analytics.Both options carry real costs. But a third option now exists: Databricks querying on-premises data directly, with no copies and no pipelines, through the open Delta Sharing protocol embedded natively in MinIO AIStor.
AIStor
AI/ML
Data Lakes & Analytics
Databricks
Why Modern AI Architecture Breaks at the Data Layer
arrow
Modern AI architecture rests on a comfortable assumption: when AI slows down, the fix is more compute or a better model. Bigger GPUs. Denser clusters. New architectures. That assumption is now costing organizations real money.
AI/ML
Storage & Infrastructure
AIStor
Cloud Infrastructure
Building a RAG Lab with AIStor and Milvus
arrow
Learn how to build a vector database lab with Milvus and AIStor
AI/ML
Benchmarking Vector Index Creation with MinIO AIStor, Milvus, and NVIDIA cuVS
arrow
106 million vectors indexed 12x faster. AIStor with NVIDIA cuVS and GPUDirect RDMA rewrites the benchmark.
AI/ML
AIStor
Performance
Integrations & Partners
Storage & Infrastructure
Long corridor in a data center with rows of illuminated servers and blue lighting reflections.
AIStor Inside NVIDIA BlueField-4: Object Data at Wire Speed
arrow
Legacy storage talks to the AI factory. AIStor lives inside it — running natively on NVIDIA BlueField-4 Vera
AIStor
AI/ML
Performance
Storage & Infrastructure
Architecture & Design Patterns
Databricks logo with stacked blocks icon on a blurred gradient background in dark pink and orange tones.
Unlocking On-Premises Data for Databricks: Secure, Zero-Copy Sharing with AIStor Table Sharing
arrow
Delta Sharing + AIStor enables zero-copy access to on-prem data from Databricks without duplication
AIStor
Architecture & Design Patterns
Data Lakes & Analytics
Integrations & Partners
Security
MinIO AIStor with NVIDIA GPUDirect® RDMA for S3-Compatible Storage: Unlocking Performance for AI Factory Workloads
arrow
GPUDirect RDMA bypasses CPU for direct GPU-to-storage S3 data paths, boosting AI factory performance
AI/ML
Logos of MinIO AIStor, GMI, and NVIDIA on a dark gradient background.
Supercharging Inference for AI Factories: KV Cache Offload as a Memory-Hierarchy Problem
arrow
KV cache offload treats GPU memory as tiered hierarchy using object storage for overflow
AI/ML
Digital illustration of federal buildings under connected data points with text on AI readiness and data infrastructure.
Federal AI Readiness: Why Data Infrastructure Matters More Than Algorithms
arrow
Federal agencies need unified data governance and infrastructure before AI algorithms
AI/ML
Text introducing MinIO ExaPOD, the reference architecture for Exascale AI, with Intel, Solidigm, Supermicro logos.
Introducing MinIO ExaPOD: The Reference Architecture for Exascale AI
arrow
Introducing ExaPOD—MinIO's 1-exabyte reference architecture with Intel, Solidigm & Supermicro for exascale AI
AI/ML
From hours to minutes: Automating cluster deployments with AIStor MCP server
arrow
Automate cluster deployments with MCP—AIStor MCP server cuts demo environment setup from hours to minutes using Claude
AI/ML
U.S. Capitol building illuminated at night with digital blue data streams extending from it.
Solving The Challenges of On-Prem Sovereign AI for Government Agencies
arrow
Solve sovereign AI challenges for government—on-prem infrastructure ensures secure, compliant AI under national control
AIStor
Case Studies & Solutions
AI/ML