How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
Hugging Face published an account of how Papers with Code, which it began reviving three months ago, implements search using three Hugging Face services: Jobs, Storage Buckets, and Inference Endpoints.
The system maintains embeddings for more than 110,000 current papers sourced from arXiv and Daily Papers. It is described as a hybrid search system that combines keyword search, provided by PostgreSQL full-text search, with dense vector search using pgvector, fused by the reciprocal rank fusion (RRF) algorithm with equal branch weights and k=60.
The architecture splits an offline corpus build from an online search service. Corpus embedding runs as Hugging Face Jobs on a GPU, with artifacts synced to a private Storage Bucket via hf-mount and mounted into an l4x1 Job, an NVIDIA L4 GPU with 24GB of VRAM. The production generation uses Qwen/Qwen3-Embedding-0.6B pinned to an exact revision, with 256-dimensional L2-normalized vectors. In a 5,000-paper pilot, the Qwen Job encoded about 75 papers per second at 1024 dimensions, and the 256-dimensional index achieved 0.9955 Recall@20 against exact search with 1.31 ms p50 and 2.21 ms p95 HNSW lookup latency.
A live query is embedded through an authenticated Inference Endpoint backed by Text Embeddings Inference (TEI). The endpoint has a maximum of one replica and can scale to zero when idle; the query client uses a one-second timeout, a circuit breaker, and no raw query text in logs. If the endpoint is scaling up, times out, returns a malformed vector, or has no concurrency available, the semantic branch is skipped and lexical results are returned.
An hourly incremental process sends up to 500 changed or missing papers per run to the same endpoint, using the document prompt in batches of 16. Buckets hold immutable run directories with manifests and SHA-256 checksums, supporting resumable Jobs, comparable experiments across model and dimension choices, and rollback before a new generation is activated. The same document embeddings also power related-paper recommendations, which need no model call at request time.
Based on reporting from the original publisher. Visit the source for full context and later updates.