Infino runs BM25, vector, hybrid, and SQL on one Parquet table

Infino is an open-source retrieval engine that runs natively on Apache Parquet files in object storage (S3, GCS, Azure Blob) or local disk. It provides four query modes in one library:

  • BM25 full-text search — 125µs p50 for 1M documents
  • Vector cosine similarity — 5ms p99 for 1M vectors
  • Hybrid search — fused BM25 + vector ranking in one SQL call
  • SQL — Apache DataFusion with search-index pushdown for filtered aggregations

The architecture stores columns and both search indexes (keyword and vector) inside each Parquet file. Tables larger than RAM work via byte-range reads from object storage. No daemon, no cluster, no lock service required.

Agent-scale retrieval demand drives the rewrite

Agent-scale development is driving demand for retrieval infrastructure that can handle many concurrent, low-latency queries against large corpora. GitHub reported 3.35B monthly pushes in 2026 — 4.9x year-over-year growth. Every agent commit needs to read context: logs, documentation, code. The traditional stack (Elasticsearch + Pinecone vectors + Neo4j graphs + ClickHouse analytics + Snowflake warehouses) becomes a cost and complexity nightmare at this scale.

Apache-2.0 crate at 0.x with production users

  • License: Apache-2.0
  • Repository: github.com/infino-ai/infino (663 commits, 112 stars, 21 forks)
  • Language: Rust with Python and Node.js bindings
  • Maturity: 0.x crate (API can still move), production use by United States Government, Turiance AI, Bazinga Labs

Pick Infino for agent retrieval at 1M+ documents

Pick Infino when:

  • Running constant agent retrieval against 1M+ documents
  • Cost sensitivity matters (10–23x cheaper than Elasticsearch/OpenSearch)
  • Want a single storage format serving search, analytics, and agents
  • Need zero operational overhead (library, not a service)

Default to Elasticsearch/OpenSearch when:

  • Need concurrent writers on the same table
  • Require mature cluster management and monitoring
  • Have existing Elasticsearch-based tooling

Single writer, 0.x API, no built-in replication

  • Single writer per table (append-only, tombstone deletes). Concurrent writers require sharding.
  • Crate is 0.x; API can still move.
  • No built-in replication or failover — relies on object storage durability.
  • Vector index recall at 0.992@10 (tested against brute-force). Not 1.0.

Source, docs, architecture, playground, pricing

  • Source: github.com/infino-ai/infino
  • Documentation: infino.ai/docs
  • Architecture: infino.ai/architecture
  • Playground: infino.ai/playground
  • Pricing methodology: infino.ai/pricing-methodology