AI & Computational Science

Static Bootstrap Placement for Encrypted Language Model Decoding

How the science connects

Language modelHomomorphic encryp…Bootstrapping

AI Insight

Researchers developed AR-HE, a system that enables language models to generate text while keeping user prompts encrypted throughout the entire process, eliminating the need for the client to communicate with the server after each generated token. The system uses a novel approach to place cryptographic bootstrapping operations strategically, reducing the time to generate a single GPT-2 token from 4715 seconds to 544 seconds on an NVIDIA H100 GPU. Key optimizations include an encrypted cache for storing previous context and a formula-based method for determining bootstrap placement without requiring search algorithms.


This advancement makes privacy-preserving AI inference more practical by allowing users to query language models without exposing sensitive information in their prompts, addressing growing concerns about data privacy in cloud-based AI services. The significant performance improvements bring fully encrypted language model inference closer to real-world feasibility, though generation times remain substantially slower than unencrypted inference.


Understand the Science

Language model 27 articles Explore Concept → Homomorphic encryption Concept coming soon Bootstrapping Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generating text with them requires a client round trip at every generated token. Keeping the loop on the server instead requires selecting and consuming a token under encryption, and placing bootstraps for a loop body that grows with the context. We build AR-HE, which runs the whole loop on the server, selects each token under encryption, retrieves its embedding, and writes it back into the encrypted state. The client sends one prompt and remains offline until the output. One rule places every bootstrap in the run, without search, so the bootstrap cost of a token is a formula in the context length that is known before the run starts. The schedule skips work whose result cannot reach the output, packs bootstraps that share an operand, and keeps the keys and values of past positions in an encrypted cache. With every optimization applied, generating a GPT-2 small token costs 544 seconds on one NVIDIA H100, down from 4715 seconds without optimization. The prompt step before it costs 4630 seconds. The cache alone takes a generated step from 11751 bootstraps to 1072. The formula predicts every step we measured, including steps of a model it was not derived from.

Source: Static Bootstrap Placement for Encrypted Language Model Decoding