DeepSeek V4.1 Currently 1500x Cheaper Than Usual
DeepSeek V4.1 Flash is at present listed at concerning 1, 500 times cheaper than DeepSeek's usual off-peak input rate on some third-party hosts.
DeepSeek V4.1 Flash is at present listed at concerning 1, 500 times cheaper than DeepSeek's usual off-peak input rate on some third-party hosts.
Article outline
- What happened
- The key numbers
- Background
- The bottom line
Key points
- DeepSeek V4.1 Flash is a sparse mixture-of-experts model published around September 10, 2026.
- Relace's OpenRouter row at present indicates $0.0001 / $0.60 per million for input and output.
- Output cost still sits near DeepSeek's off-peak $0.60 on Relace, so chatty or long-answer jobs will not see the same 1, 500x savings.
- DeepSeek's published off-peak schedule for V4.1 Flash is $0.15 per million input tokens and $0.60 per million output tokens, with peak rates doubling both figures.
- That gap is what sparked a Reddit thread on r/DeepSeek titled around Relace and Open Inference's "shockingly low" listed rates.
That gap is what sparked a Reddit thread on r/DeepSeek titled around Relace and Open Inference's "shockingly low" listed rates. Open Inference's OpenRouter listing sits near $0.00011 per million input tokens and $0.36 per million output tokens.
In practice, the 1, 500x figure applies to input pricing only. Relace still lists output at $0.60 per million tokens. It matches DeepSeek's own off-peak output rate. So the bargain is on the input side of those provider listings, not a full rewrite of every token cost. How The Listed Rates Compare.
DeepSeek's published off-peak schedule for V4.1 Flash is $0.15 per million input tokens and $0.60 per million output tokens, with peak rates doubling both figures. Cache hits drop input cost further on DeepSeek's own API.
Relace's OpenRouter row at present indicates $0.0001 / $0.60 per million for input and output. Open Inference demonstrates concerning $0.00011 / $0.36. Against DeepSeek's $0.15 off-peak input rate, Relace's $0.0001 input list cost is roughly 1, 500 times lower.
Those figures are provider list rates on aggregators such as OpenRouter. Listed rates can change, and real bills additionally depend on cache hits, routing, latency, and how numerous output tokens a job uses. What DeepSeek V4.1 Flash Is.
DeepSeek V4.1 Flash is a sparse mixture-of-experts model published around September 10, 2026. It is the first model built on DeepSeek's Causal Encoder-Decoder architecture. The firm describes a 552 billion parameter backbone that activates concerning 8 billion parameters on input and 16 billion on output.
Notably, the model is aimed at coding, terminal work, computer-use agents, and long-context jobs. OpenRouter lists a 1 million token context window. DeepSeek additionally states compressed KV caching cuts cache memory versus the previous Flash generation. It matters for agent workflows that reuse long prompts.
ProPakistani earlier covered the DeepSeek V4.1 Flash launch, including its lower official pricing tier and native vision backing. The new attention is less regarding another official cut and more concerning how cheap some third-party endpoints are advertising the same model. Why Developers Are Watching The Listings.
For high-volume input workloads, a move from $0.15 to $0.0001 per million tokens would be dramatic if the endpoint stays stable and the output quality holds up. Output cost still sits near DeepSeek's off-peak $0.60 on Relace, so chatty or long-answer jobs will not see the same 1, 500x savings.
Though recent OpenRouter snapshots have additionally shown weaker uptime and slower throughput on that route compared with Relace, open Inference's lower output list cost ($0.36) is a separate pull. Developers comparing hosts still need to weigh cost against reliability and speed.
DeepSeek has additionally been covered for earlier pricing moves, including when the business raised rates by up to 4x. The Relace and Open Inference listings sit at the opposite extreme: third-party rows that undercut DeepSeek's own input sticker by orders of magnitude.
For now, the practical takeaway is narrow and checkable. While output pricing stays much closer to normal, on current OpenRouter listings, Relace and Open Inference advertise DeepSeek V4.1 Flash input rates near $0.0001 per million tokens, regarding 1, 500 times below DeepSeek's usual $0.15 off-peak input rate. Source: r/DeepSeek on Reddit; OpenRouter. Stay Connected with ProPakistani.
Obtain the latest tech news, telecom insights, and product launches wherever you prefer. Follow on Google Discover. Follow on Google News Join WhatsApp. See more ProPakistani stories in Google Search and Top Stories.
For now, deepSeek V4.1 Currently 1500x Cheaper Than Usual remains the part of the story worth watching, and further updates are likely as more details are confirmed.



