Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

AWS Compares Bedrock Vector Stores

AWS Compares Bedrock Vector Stores

AWS ML Blog·Friday, September 18, 2026
  • •AWS compares three customer-managed vector stores for Amazon Bedrock Knowledge Bases RAG deployments
  • •OpenSearch product-search benchmark indexed 1,215,851 products and sampled 5,000 judged queries
  • •512-float matched 1024-float NDCG while cutting index size from 5.34 GiB to 2.79 GiB
  • •AWS compares three customer-managed vector stores for Amazon Bedrock Knowledge Bases RAG deployments
  • •OpenSearch product-search benchmark indexed 1,215,851 products and sampled 5,000 judged queries
  • •512-float matched 1024-float NDCG while cutting index size from 5.34 GiB to 2.79 GiB
  • •AWS compares three customer-managed vector stores for Amazon Bedrock Knowledge Bases RAG deployments
  • •OpenSearch product-search benchmark indexed 1,215,851 products and sampled 5,000 judged queries
  • •512-float matched 1024-float NDCG while cutting index size from 5.34 GiB to 2.79 GiB
  • •AWS compares three customer-managed vector stores for Amazon Bedrock Knowledge Bases RAG deployments
  • •OpenSearch product-search benchmark indexed 1,215,851 products and sampled 5,000 judged queries
  • •512-float matched 1024-float NDCG while cutting index size from 5.34 GiB to 2.79 GiB

AWS published guidance on September 17, 2026, for choosing a customer-managed vector store in Amazon Bedrock Knowledge Bases when building Retrieval Augmented Generation, or RAG (retrieving relevant text before generation), systems. The post compares three supported backends: Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors, and says the selection affects performance and cost across different RAG use cases.

Amazon Bedrock Knowledge Bases offers a fully managed option and a customer-managed option. In the customer-managed path, Amazon OpenSearch Service is described as a high-speed option with managed cluster and serverless deployments, k-NN search, and hybrid search that combines lexical and vector methods. Amazon Aurora PostgreSQL with pgvector adds vector similarity search to Aurora and supports IVFFlat and HNSW indexing, L2, cosine, and inner product distance metrics, and vectors up to 2,000 dimensions in single precision. Amazon S3 Vectors is described as native vector support in Amazon S3, with sub-second similarity-search performance and vector storage cost reductions of up to 90 percent compared with traditional vector databases.

The post explains RAG as a system where documents are pre-processed, chunked, converted into vector embeddings, and stored in a vector database. When a user submits a query, the query is also converted into a vector embedding, and the system retrieves the top n chunks, such as the top five most similar chunks, before passing them to the LLM as added context. The vector database stores embeddings in a vector index (structure for fast similarity search), enabling high-dimensional semantic search rather than simple keyword matching.

For product catalog search, the article identifies Amazon OpenSearch Serverless as the best fit because ecommerce search needs natural language understanding, low latency, and the ability to handle thousands of concurrent queries during peak shopping periods. It cites hybrid search, low-milliseconds query latency, complex filtering, aggregations for faceted navigation such as price, brand, or color, and distance metrics such as cosine similarity and Euclidean distance.

The OpenSearch benchmarks used Serverless Classic collections, not Amazon OpenSearch Serverless NextGen collections, which became generally available in May 2026 and were not yet compatible with the Amazon Bedrock Knowledge Bases Retrieve API. NextGen removes engine and mode parameters from index mappings, defaults to 32× compression with GPU-accelerated index builds, and supports scale-to-zero. The article notes that Managed Clusters provide additional tuning options, including auto-optimize, GPU-accelerated indexing, and configurable instance sizing.

The benchmark used Amazon’s Shopping Queries Data Set, or ESCI, with 1,215,851 unique US products, approximately 1,140 characters median per product, and 97,345 judged queries. The test sampled 5,000 queries with approximately 19 judged products per query and approximately 17 relevant products, then indexed all 1,215,851 product descriptions. Measurements included NDCG@10, p50, p95, and p99 latency at concurrency 1 and 10, and ANN in-memory index size.

The experiment tested embedding dimensions of 1024, 512, and 256, with float and binary data types, making six configurations, plus 1024-float in on_disk mode at compression_level: 32x, for seven total configurations. All indexes used FAISS with HNSW, ef_construction=128, m=24, l2 distance for float embeddings, and hamming distance for binary embeddings. Each index was ingested with all 1.22M documents, warmed until latency stabilized, measured, deleted, and followed by a 15-minute cool down.

Results showed 512-float was statistically indistinguishable from the 1024-float baseline, with NDCG 0.3628 vs 0.3627 and p = 0.87, while cutting index size to 2.79 GiB from 5.34 GiB and lowering p50 latency to 25 ms from 31 ms. Dropping to 256 dimensions produced a 4.4 percent quality loss. Binary embeddings at 1024 dimensions reduced index size by 13.4×, to 0.40 GiB from 5.34 GiB, with a 5.2 percent NDCG loss and 22 ms vs 31 ms p50 latency. At 256 dimensions, the quality cost rose to 28.3 percent and the size saving was 5.0×. On_disk 32× at 1024 dimensions reached NDCG 0.3610, or −0.5 percent vs baseline, with 0.40 GiB index size and about 3× higher p50 latency, 99 ms vs 31 ms.

AWS published guidance on September 17, 2026, for choosing a customer-managed vector store in Amazon Bedrock Knowledge Bases when building Retrieval Augmented Generation, or RAG (retrieving relevant text before generation), systems. The post compares three supported backends: Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors, and says the selection affects performance and cost across different RAG use cases.

Amazon Bedrock Knowledge Bases offers a fully managed option and a customer-managed option. In the customer-managed path, Amazon OpenSearch Service is described as a high-speed option with managed cluster and serverless deployments, k-NN search, and hybrid search that combines lexical and vector methods. Amazon Aurora PostgreSQL with pgvector adds vector similarity search to Aurora and supports IVFFlat and HNSW indexing, L2, cosine, and inner product distance metrics, and vectors up to 2,000 dimensions in single precision. Amazon S3 Vectors is described as native vector support in Amazon S3, with sub-second similarity-search performance and vector storage cost reductions of up to 90 percent compared with traditional vector databases.

The post explains RAG as a system where documents are pre-processed, chunked, converted into vector embeddings, and stored in a vector database. When a user submits a query, the query is also converted into a vector embedding, and the system retrieves the top n chunks, such as the top five most similar chunks, before passing them to the LLM as added context. The vector database stores embeddings in a vector index (structure for fast similarity search), enabling high-dimensional semantic search rather than simple keyword matching.

For product catalog search, the article identifies Amazon OpenSearch Serverless as the best fit because ecommerce search needs natural language understanding, low latency, and the ability to handle thousands of concurrent queries during peak shopping periods. It cites hybrid search, low-milliseconds query latency, complex filtering, aggregations for faceted navigation such as price, brand, or color, and distance metrics such as cosine similarity and Euclidean distance.

The OpenSearch benchmarks used Serverless Classic collections, not Amazon OpenSearch Serverless NextGen collections, which became generally available in May 2026 and were not yet compatible with the Amazon Bedrock Knowledge Bases Retrieve API. NextGen removes engine and mode parameters from index mappings, defaults to 32× compression with GPU-accelerated index builds, and supports scale-to-zero. The article notes that Managed Clusters provide additional tuning options, including auto-optimize, GPU-accelerated indexing, and configurable instance sizing.

The benchmark used Amazon’s Shopping Queries Data Set, or ESCI, with 1,215,851 unique US products, approximately 1,140 characters median per product, and 97,345 judged queries. The test sampled 5,000 queries with approximately 19 judged products per query and approximately 17 relevant products, then indexed all 1,215,851 product descriptions. Measurements included NDCG@10, p50, p95, and p99 latency at concurrency 1 and 10, and ANN in-memory index size.

The experiment tested embedding dimensions of 1024, 512, and 256, with float and binary data types, making six configurations, plus 1024-float in on_disk mode at compression_level: 32x, for seven total configurations. All indexes used FAISS with HNSW, ef_construction=128, m=24, l2 distance for float embeddings, and hamming distance for binary embeddings. Each index was ingested with all 1.22M documents, warmed until latency stabilized, measured, deleted, and followed by a 15-minute cool down.

Results showed 512-float was statistically indistinguishable from the 1024-float baseline, with NDCG 0.3628 vs 0.3627 and p = 0.87, while cutting index size to 2.79 GiB from 5.34 GiB and lowering p50 latency to 25 ms from 31 ms. Dropping to 256 dimensions produced a 4.4 percent quality loss. Binary embeddings at 1024 dimensions reduced index size by 13.4×, to 0.40 GiB from 5.34 GiB, with a 5.2 percent NDCG loss and 22 ms vs 31 ms p50 latency. At 256 dimensions, the quality cost rose to 28.3 percent and the size saving was 5.0×. On_disk 32× at 1024 dimensions reached NDCG 0.3610, or −0.5 percent vs baseline, with 0.40 GiB index size and about 3× higher p50 latency, 99 ms vs 31 ms.

Read original (English)·Sep 17, 2026
Infra#amazon bedrock#knowledge bases#vector store#rag#opensearch#pgvector#s3 vectors#hnsw#faiss#ndcg