How Milvus Handles Billion-Scale Vectors with Cost-Efficient Distributed Design

Zilliz
254 views June 30, 2025

At Zilliz, we didn’t take shortcuts when building for billion-scale vector search — we leaned into hard engineering trade-offs to give you real cost and performance wins. In this week's AWS GenAI livestream, Jiang Chen from Zilliz broke down why and how we embraced engineering complexity to build Milvus' distributed architecture – to create a vector database that can scale to billions while maintaining cost-efficiency. Here's what he shared: 1. Distributed Architecture – we took extra engineering complexity to make Milvus scale horizontally so that our users can focus on application building without worrying about the infrastructure scalability. 2. Affordable Search at Billion Scale – Serving all vector index from memory offers the best speed — but it comes at a high cost. Milvus supports quantization, on-disk index type, and tiered storage design to cache data in expensive RAM, NVMe SSDs, and persisted in S3, so you can dial down cost to your latency expectations without blowing your infra budget. 3. Introducing Vector Lake – purpose-built for offline & interactive analytics. Keep petabytes of embeddings in S3-backed cold storage, only spin up the query engines to serve the ad-hoc load. This will make trillion scale offline search work economically.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close