A well-designed token, a real customer, and one binary catalyst still to come.
Voorhees needed uncensored models for Venice, so he funded a small lab called Dolphin in 2023. Two years later the lab has three million users routed through its models on Venice, five million monthly Hugging Face downloads, and a token whose contract spends every dollar of network revenue buying POD on the open market. The mechanism is coded. The revenue is not, because the public API has not opened. Whether the door opens in the next four to eight weeks is the whole trade.
Dolphin built its customer before it built its network. That single fact is what separates it from the dozens of decentralized inference projects incentivising GPUs and waiting for demand to arrive.
Erik Voorhees was building a consumer AI product, Venice, that would sell privacy and uncensored responses. To do that he needed models that would answer the question when asked, and most of the open frontier would not. So he funded a small lab that spent two years quietly fine-tuning uncensored versions of Llama, Mistral, Qwen, and Gemma.
By the time that lab, Dolphin, issued a token, its models were the default uncensored layer on Venice for three million users, and its open releases pulled five million downloads a month on Hugging Face. Only then did they stand up a network to serve them.
The current network state reflects that ordering. Peak stress testing ran 14,000 GPUs through 1.4 trillion tokens in two weeks. Supply is intentionally suppressed right now to around 800 GPUs at $0.10 per million tokens in rewards, held low while the team finishes anti-cheat work and prepares to migrate to Qwen 3.6 27B ahead of the API launch. That capacity is a demonstrated ceiling, not a running rate. The workload is still a coding dataset the network is generating for itself, because the public API has not opened. Nodes earn POD emissions from treasury; there is no external revenue to buy back against yet. That is the timing risk in one sentence.
Two things separate POD from every other DePIN inference token I have looked at this year. Verification that actually holds up. Value capture that actually flows.
The hard problem in decentralized inference is knowing that the stranger you paid ran the model they said, at the quality they claimed, and not a smaller quantised version they had in a folder. Every prior attempt has hit the same wall. File-hash checks fail because nodes can swap models in memory after the check. Prompt re-execution works but leaks the prompt and doesn’t scale.
Dolphin’s answer is live-weight proofs. Sample the tensors actually resident in the runtime at the moment a response is produced, compare against the approved model’s manifest. Overhead is around 0.1% of a full re-inference. Cheap enough to run on every response. It verifies what is loaded, not what is emitted, so it extends to image, audio, and video where token-based verification breaks down.
Which matters, because the cost advantage is largest in exactly those modalities. A 4090 is 12 to 18 times more efficient per dollar than an H100 for audio, and 5 to 8 times for image and video. Those are also the highest-margin modalities in AI. ElevenLabs alone is reportedly near $500M ARR on audio. And the video read has just shifted: MiniMax H3, released in July, matches Sedance quality on a single RTX 6000 with audio built in. That moves the video tier from speculative to plausible on the roadmap.
100% of network revenue automatically buys POD on the open market. No discretionary schedule, no per-epoch team decision. Coded into the contract.
The bought-back POD doesn’t burn. It splits between the treasury (which pays node emissions) and the xPOD staking vault (as auto-compounding yield to stakers). Governance sets the split ratio, and the ratio is still to be decided. That is the softest part of the design.
For a passive holder this is closer to negative-carry than to a share of cash flows. To capture the mechanism, you stake into xPOD, accept a three-month cooldown to exit, and take yield as auto-compounded POD.
| Line | Per 1M tokens | |
|---|---|---|
| Cheapest OpenRouter comparable | $1.00 | |
| Dolphin list price | $0.70 | |
| Paid to node operator | $0.50 | |
| Network spread → POD buyback | $0.20 | |
Undercutting the cheapest centralised provider by 30% while still extracting $0.20 of buyback per million tokens is a real edge. On the demand side, Qwen 3.6 27B is claimed at 82 to 88% of Opus 4.7 on coding tasks. Opus runs $25 per million output tokens. Dolphin at $0.70 is about 3% of the frontier price for mid-eighties frontier quality.
OpenRouter routes about a billion dollars a year in inference. Even 1% capture is $10M gross and $2.9M of buyback. Against a $15M market cap that is a 19% annualized bid on the token. Against $172M FDV it is 1.7%. Reflexivity works on the near-term price and competes with fully diluted supply long-term. Fine, provided you are honest about the timeframe you are trading.
That math is the wholesale API. The team is building two adjacent surfaces on top of it. Flipper is a coding harness targeting Claude Code and Codex on cost, with an MVP in about four weeks, using 27B subagents planned by Qwen Max. Sticky developer relationships convert better than API rentals, and this is the more defensible surface long-term. Fusion is an intelligent router: most requests to 27B, complex ones to a larger on-network model (Deepseek Flash, 300B parameters), or out to Venice for frontier quality, paid in API credits from a GLM 700B fine-tune Dolphin is finishing for Venice right now. None of this changes the buyback mechanism. All of it feeds the same contract.
| Bucket | POD | Terms |
|---|---|---|
| Treasury | 289.6M / 57.9% | Protocol-controlled. Drips out via node emissions at $0.50/M tokens. |
| Seed | 117.6M / 23.5% | $886k raised Jun 2024 at $0.0075. Cliff cleared Jul 1, 2026. |
| Team | 50M / 10.0% | 2yr cliff + 10yr linear from Jun 2026. Fully vested Jun 2036. |
| LP + OTC | 47M / 9.5% | Uniswap v4 pool plus three staked OTC deals. |
Nobody signs a twelve-year vest without believing in the outcome. The team also hard-committed to no equity. The Cayman service company that will accept fiat and list on OpenRouter was formed in mid-July, with a director appointed. It remits 99% of collected revenue to the protocol and keeps 1% as an administrative margin. The token remains the only value-capture instrument. That combination, more than any single mechanism, is what makes this a real bet rather than a fair-launch fantasy.
The natural comparison is Akash, the category’s most established name. Akash is a real project with real traction: meaningful compute spend, a Burn-Mint Equilibrium settlement stablecoin, and through AkashML reportedly serving over a billion tokens a day on OpenRouter. It even added a consumer-GPU beta to reach the same hardware pool. Nothing about Akash is weak. The gap is structural.
Akash sells raw compute and sources demand externally. Dolphin’s distribution runs through Venice, five million monthly Hugging Face downloads, and integrations with Cursor, Cline, and Roo. Akash has no inference-integrity primitive comparable to live-weight proofs, and session-rental supply cannot reach the idle-gamer-GPU pool peer-to-pool absorbs by design.
The tell is on the pricing side. AkashML inference is reportedly priced higher than Alibaba’s own first-party Qwen API. A decentralized network is supposed to undercut the centralised incumbent, not charge more than the model’s creator. That is exactly the opening a leaner consumer-GPU network is built to exploit.
Akash trades at $210M. Dolphin trades at $15M. If Dolphin lands the API and the OpenRouter listing, closing to Akash’s zone is 14x. If it doesn’t, the gap is priced right, and probably widens the wrong way.
The May draft passed three of six cleanly and left three pending on token disclosure. TGE has happened. The doc is public. The picture now:
| Filter | Verdict | Note |
|---|---|---|
| Working product | Pass | 5M monthly downloads, 3M users through Venice, network live. |
| Mechanical demand | Pass on design | Buyback is coded. Missing input is external revenue. |
| Incumbents can’t capture | Pass | SLA-based providers can’t route to consumer hardware. |
| Founder alignment | Pass | 12-year vest, no equity, non-profit fiat entity if ever needed. |
| Asymmetric entry | Pass on MC | $15M MC clean. $172M FDV against Akash’s $210M less so. |
| Reflexivity favours holders | Pass on design | Cash-flow accrual to xPOD. Loop hasn’t started spinning. |
Five clean passes and one pending on activation. The framework that had three unknowns in May now hangs on one thing: whether the API opens.
The strongest short on POD isn’t complicated. It runs like this. The API never opens on a schedule that matters, because Dolphin can’t ship a public inference product that works at scale, because the team is optimising for headlines and dataset generation over the harder work of onboarding paying developers. Meanwhile 14,000 GPUs burn treasury POD for a coding dataset that maybe sells and maybe doesn’t. The Qwen benchmarks turn out to be Qwen’s numbers, not the community’s. Qwen 27B lands at 68% of Opus 4.7 rather than 88%, and the “3% of frontier price for mid-eighties quality” pitch collapses to “cheaper than a model buyers don’t want.” Venice plateaus at three million users because uncensored is a niche, not a market. Prime Intellect ships equivalent verification on a bigger network before Dolphin closes its gap. In that world POD is a beautifully-designed mechanism that never gets to prove itself, and the emissions clock runs out before the demand clock arrives.
None of the pieces is the base case. But you don’t need all of them. Two of the five is enough to break the trade, and the sizing has to survive that.
Starter position, 0.5 to 1% of the AI-crypto sleeve, at $0.30 to $0.35. Stake into xPOD immediately. Treat as illiquid for three months minimum.
If you’re already holding VVV or sVVV as I am, you’re already exposed to Dolphin whether you meant to be or not. Venice’s success feeds Dolphin’s revenue funnel, and Dolphin’s uncensored models are what Venice sells. A meaningful POD position on top of existing Venice exposure turns three positions into one bet expressed three ways. That is real concentration. Size the starter accordingly, and do not add before you decide whether the concentration is a feature or a bug.
Scale up if any of these land: V2 worker release ships, the public API opens, an OpenRouter listing goes live near the $0.70 mark, first on-chain buyback executions show visible volume, governance publishes the buyback-to-xPOD split at 50% or higher.
Cut the position if the API slips past October with no update, or first-month OpenRouter revenue is under $100k, or independent Qwen 3.6 27B benchmarks come in under 75% of Opus 5, or the buyback contract shows negligible activity 60 days after launch, or Venice reduces routing volume to Dolphin models.
Four things I don’t know. When the API opens (soon has been the answer for six weeks). How the buyback split gets set. Whether the Qwen benchmarks survive independent evaluation. Whether Venice keeps growing past three million.
None of that changes the sizing. It changes the willingness to hold. If the answers come back wrong, the position closes without argument. If they come back right, this looks obvious in retrospect and probably too small.
Everything in this memo resolves to one gate. The public API opens on a schedule that beats the emissions clock, or it does not. Three signals to watch: the V2 worker release, an OpenRouter listing near $0.70, and the first visible buyback executions from the treasury address on Base. Those three resolve the trade.