
Ethereum co-founder Vitalik Buterin has tested a privacy-focused AI setup that uses a local model, zkAPI and Tor to generate personalized diet and exercise recommendations while limiting personal information sent to remote models.
Summary
- Vitalik Buterin is testing private AI health advice using local Qwen, zkAPI, and Tor routing.
- His local model rewrites prompts before remote models receive limited health and travel data remotely.
- zkAPI separates payment identity from model requests, while Tor is used to mask IP information.
- Buterin said Tor latency remains 10–100 times higher, making request-by-request unlinking inefficient in current tests.
- Qwen3.8-Flash-Next runs around 20–30 TPS locally, while Buterin wants speeds above 100 TPS for comfort.
Buterin said on Oct. 4 that the self-experiment uses his health and travel information locally, while more powerful remote models handle selected questions that require stronger reasoning or knowledge.
The setup uses Alibaba’s Qwen3.8-Flash-Next as the local model. Buterin said the local system decides what information a remote model needs and rewrites requests before sending them, reducing the chance that personal details or his writing style reveal his identity.
Vitalik Buterin uses three layers to separate his identity
Buterin described the design as a three-layer privacy setup covering the content of requests, payment information and internet traffic. The local Qwen model handles the first layer by constructing queries itself instead of sending his original wording and full personal context to remote AI systems.
The second layer uses zkAPI to separate payments from individual AI requests. The Ethereum Foundation introduced zkAPI on Oct. 1, describing it as a system that lets users pay for metered APIs without linking individual requests to their identity. The project was built by the Open Anonymity Project in collaboration with the Ethereum Foundation and is running on Ethereum mainnet.
Under zkAPI, a user funds a private balance and later proves that enough funds are available without showing which deposit is paying for a particular request. The service handling payment does not need the user’s prompt, while the AI provider receives the prompt without learning the billing identity tied to the deposit.
Tor provides the third layer by hiding the user’s normal IP address from the services receiving network requests. Buterin wrote that all three protections are needed because hiding payment information alone does not prevent an AI provider from learning details through prompt content or network metadata.
“You need all three,” Buterin said.
zkAPI does not hide everything sent to an AI model
The privacy setup does not prevent remote AI providers from reading information deliberately included in a prompt. Official zkAPI documentation states that the upstream provider still sees prompts, while network and timing information can remain observable outside the zero-knowledge proof system.
The Ethereum Foundation made the same distinction when it launched zkAPI. Its Oct. 1 explanation said the payment system hides the link between a user and a request, but content privacy and network anonymity require separate protections. Reused personal details, writing patterns, conversation history or documents can still allow sessions to be connected.
Buterin’s local model is intended to reduce that content exposure. A skill file instructs the model when to use a remote system and how to construct a request that contains less identifying information. His personal health and travel records remain available to the local system, while the remote model receives only the portion selected for a particular task.
Buterin said the setup produced diet and exercise recommendations and that information returned by frontier models improved the results. He did not publish the underlying health records, the detailed recommendations or an independent evaluation of their accuracy.
The experiment fits with his earlier focus on privacy as AI systems handle more personal information. As previously reported in crypto.news coverage of Buterin’s privacy concerns, he argued in April 2025 that growing AI capabilities and centralized data collection increased the need for stronger privacy tools.
Tor support has reached the zkAPI codebase
Buterin linked to a new change in the Ethereum zkAPI repository that adds Tor-routed client support. GitHub shows pull request #16 as open as of Oct. 4, with one commit proposing changes across seven files. It has not yet been merged into the project’s main branch.
The proposed code creates a fresh temporary Tor client when the zkAPI daemon starts. The script uses a new data directory and Tor connection, while another command can restart the service for a fresh network identity before a new single request or conversation begins.
The patch changes several network timeouts because requests routed through Tor can take longer. One model-list timeout rises from one minute to three minutes, while other request limits increase from 15 seconds to 60 seconds and from five seconds to 30 seconds.
A separate Tor client script included in the proposal says a fresh server is created for a single request or the start of a new conversation. Continued messages within the same conversation keep the existing server running, meaning they do not automatically receive a new Tor identity for every message.
Tor latency and local AI speed remain problems
Buterin identified Tor as one of the weakest parts of the current experiment. He said Tor was not designed for the type of request-by-request unlinking he wants, where separate AI calls would ideally be difficult to associate with one another.
In his testing, Tor produced latency roughly 10 to 100 times higher than what he considered desirable. The GitHub changes increasing several timeout limits are consistent with slower network requests being expected when the zkAPI client is routed through Tor.
The local model presents another performance limit. Buterin said Qwen3.8-Flash-Next was running at approximately 20 to 30 tokens per second in his setup, but he believed local inference would only begin to feel fast at more than 100 tokens per second.
Alibaba’s Qwen team released Qwen3.8-Flash-Next on Aug. 26. The official repository describes it as an open-weight foundation model that can run through local inference frameworks, including deployments using vLLM and SGLang.
Buterin had already been experimenting with local Qwen models before the latest privacy test. His current setup goes a step further by letting the local model act as an intermediary between private files and remote AI systems instead of keeping every task entirely on the user’s device.
Privacy has remained part of his Ethereum work as well. In related coverage,crypto.news reported on Ethereum’s updated roadmap in August, which included stronger protocol privacy alongside work on quantum resistance and native rollups.
Buterin said the request-writing rules in his current experiment still need improvement because removing more personal context can reduce the usefulness of remote models. He described the limit directly: “the more careful you are” with information sent remotely, the less assistance the remote model can provide.
The Ethereum Foundation’s zkAPI documentation makes a similar technical distinction. The payment layer can break the connection between a funded balance and individual API use, but it cannot remove identifying information that a user or local agent places inside the prompt itself.