Abstract

Advanced SmartNICs known as Data Processing Units (DPU) enable in-network data analytics, being equipped with standard processors and accelerators capable of doing custom computation on-NIC. These SmartNICs are advantageous because the host CPU can delegate simpler data analytics tasks, like filtering, to the NIC where the data first arrives. Only data needing further refinement must be passed on to the host CPU. Our benchmarks focus on NVIDIA's commercially available Bluefield-3 DPU. Similarity search, particularly Approximate Nearest Neighbor (ANN) search, is an important domain with wide usage across numerous applications. We explore ANN search on SmartNICs, providing insight into the performance of various ANN algorithms on the DPU when compared to a standard CPU device. We evaluate common ANN methods such as LSH, PQ, IVFPQ, and HNSW via the FAISS (Facebook AI Similarity Search) library on a host x86 CPU and a DPU using GloVe-200 vectors. In short, HNSW delivers the best latency-recall tradeoff on both platforms; LSH suffers the steepest recall degradation. The host/DPU performance gap varies by algorithm, reflecting the devices' different strengths.

Department(s)

Computer Science

Publication Status

Free Access

Keywords and Phrases

Approximate Nearest Neighbor; BlueField-3 DPU; FAISS; GloVe; HNSW; LSH; Product Quantization; Recall@k; Vector Search

Document Type

Article - Conference proceedings

Document Version

Citation

File Type

text

Language(s)

English

Rights

© 2026 The Author(s), All rights reserved.

Creative Commons Licensing

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Publication Date

13 Jul 2026

Share

 
COinS