Infino runs BM25, vector, hybrid, and SQL on one Parquet table
Infino is an open-source retrieval engine that runs natively on Apache Parquet files in object storage (S3, GCS, Azure Blob) or local disk. It provides four query modes in one library:
- BM25 full-text search — 125µs p50 for 1M documents
- Vector cosine similarity — 5ms p99 for 1M vectors
- Hybrid search — fused BM25 + vector ranking in one SQL call
- SQL — Apache DataFusion with search-index pushdown for filtered aggregations
The architecture stores columns and both search indexes (keyword and vector) inside each Parquet file. Tables larger than RAM work via byte-range reads from object storage. No daemon, no cluster, no lock service required.
Agent-scale retrieval demand drives the rewrite
Agent-scale development is driving demand for retrieval infrastructure that can handle many concurrent, low-latency queries against large corpora. GitHub reported 3.35B monthly pushes in 2026 — 4.9x year-over-year growth. Every agent commit needs to read context: logs, documentation, code. The traditional stack (Elasticsearch + Pinecone vectors + Neo4j graphs + ClickHouse analytics + Snowflake warehouses) becomes a cost and complexity nightmare at this scale.
Apache-2.0 crate at 0.x with production users
- License: Apache-2.0
- Repository: github.com/infino-ai/infino (663 commits, 112 stars, 21 forks)
- Language: Rust with Python and Node.js bindings
- Maturity: 0.x crate (API can still move), production use by United States Government, Turiance AI, Bazinga Labs
Pick Infino for agent retrieval at 1M+ documents
Pick Infino when:
- Running constant agent retrieval against 1M+ documents
- Cost sensitivity matters (10–23x cheaper than Elasticsearch/OpenSearch)
- Want a single storage format serving search, analytics, and agents
- Need zero operational overhead (library, not a service)
Default to Elasticsearch/OpenSearch when:
- Need concurrent writers on the same table
- Require mature cluster management and monitoring
- Have existing Elasticsearch-based tooling
Single writer, 0.x API, no built-in replication
- Single writer per table (append-only, tombstone deletes). Concurrent writers require sharding.
- Crate is 0.x; API can still move.
- No built-in replication or failover — relies on object storage durability.
- Vector index recall at 0.992@10 (tested against brute-force). Not 1.0.
Source, docs, architecture, playground, pricing
- Source: github.com/infino-ai/infino
- Documentation: infino.ai/docs
- Architecture: infino.ai/architecture
- Playground: infino.ai/playground
- Pricing methodology: infino.ai/pricing-methodology
