Skip to content

DeepSeek's New Open Weights Are 167 GB. Its Pricing Page Also Warns of a 2x Peak Surcharge.

August 3, 2026. DeepSeek pushed a new build of its Flash model to Hugging Face on Friday under an MIT licence, and it is now the version its own API serves. The interesting number is not a benchmark. It is the size of the repository, and a two line footnote on DeepSeek's pricing page that nobody has covered.

What is actually on the servers

  1. The weights are live and permissively licensed. The repository deepseek-ai/DeepSeek-V4-Flash-0731 was created on July 31, 2026 at 07:30 UTC and last modified the same day at 12:02 UTC. Its card declares an MIT licence, the most permissive terms any frontier scale open weight release has carried this year.
  2. It is 48 shards and about 167 GB of storage. The repository's published storage figure is roughly 167 gigabytes across 48 safetensors shards, with the manifest counting about 304 billion parameters, the large majority of them held as 8 bit integers.
  3. It is the version the API now serves. DeepSeek's models and pricing page lists DeepSeek-V4-Flash-0731 as the model version behind deepseek-v4-flash, with a 1 million token context window and a maximum output of 384,000 tokens.
  4. The price did not move. deepseek-v4-flash remains $0.14 per million input tokens on a cache miss, $0.0028 per million on a cache hit, and $0.28 per million output tokens.
  5. A 2x peak surcharge is coming, with no date. A footnote on the same page states that the API will soon adopt peak and off peak pricing, that peak prices will be double the regular prices across all billing items, and that the effective date is subject to a future announcement. Peak hours are given as 9:00 to 12:00 and 14:00 to 18:00 Beijing time.
  6. DeepSeek has not written a release note. The company's public change log still shows April 24, 2026 as its most recent entry, so the pricing page and the repository itself are currently the only primary confirmation that this build shipped.

167 GB is a different conversation to 1.4 TB

We wrote on July 30 that Kimi K3's open weights were real and largely theoretical, because 1.4 terabytes and 64 or more accelerators put self hosting out of reach of every business we work with. This release sits in a different bracket. At roughly 167 gigabytes it is about eight times smaller, which is the difference between a model you can only rent and one a single well specified server can hold.

That does not make self hosting the right answer. Renting inference at fourteen cents per million input tokens is cheaper than owning hardware for almost everybody reading this. What it changes is your negotiating position, because a model you could run is a ceiling on the price of the model you rent.

The footnote is worth more than the benchmark

The undated 2x peak surcharge deserves far more attention than it has had. Doubling every billing item for seven hours a day is a large change to publish in a footnote, and the timing detail matters more than it first appears.

Peak is defined in Beijing time, UTC plus 8. That puts the two windows at roughly 01:00 to 04:00 and 06:00 to 10:00 UTC. A United States business running batch work inside its own office hours falls outside both windows entirely. A European team starting at 08:00 local time sits inside the second one.

The other number worth staring at is the cache hit price. At $0.0028 against $0.14, a cached input token costs one fiftieth of an uncached one. For any agent that reuses a long system prompt, prompt structure is a bigger cost lever than model choice, and unlike vendor pricing it is entirely within your control.

What it means for operators

  1. Do not switch models on a benchmark you cannot audit. DeepSeek has published no release note for this build. Test it on your own task with your own inputs before moving anything that matters.
  2. Find out when your provider's peak hours are. If DeepSeek sits anywhere in your stack, the surcharge is announced but undated. Move batch and enrichment jobs outside 01:00 to 04:00 and 06:00 to 10:00 UTC now, while being early costs you nothing.
  3. Restructure prompts before you shop for a cheaper model. A stable prefix that caches beats a 20 percent cheaper per token rate by a wide margin when the cache ratio is 50 to 1.
  4. Treat licence terms as a procurement question. MIT is unusually permissive for a model at this scale, and it is a legitimate reason to prefer one vendor over another when the capability is close.

We do this arithmetic before writing any code, whether the job is launching a SaaS product on top of a model or wiring one into a workflow that already exists. If you want it done properly, hiring an AI engineer for a week is usually cheaper than a quarter of the wrong inference bill.

Want your AI inference bill modelled before you commit?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

Yes. The Hugging Face repository deepseek-ai/DeepSeek-V4-Flash-0731 was created on July 31, 2026 at 07:30 UTC under an MIT licence, and DeepSeek's own models and pricing page now lists DeepSeek-V4-Flash-0731 as the model version served behind deepseek-v4-flash. DeepSeek's public change log has not been updated since April 24, 2026, so there is no official release note for it.

It is far more realistic than recent open weight releases. The repository publishes a storage figure of roughly 167 gigabytes across 48 shards, around eight times smaller than Kimi K3. For most businesses renting inference is still cheaper than owning hardware, at $0.14 per million input tokens on a cache miss, but the option existing is what keeps the rented price honest.

DeepSeek's pricing page states that the API will soon adopt peak and off peak pricing, with peak prices at 2x the regular prices across all billing items, and peak hours of 9:00 to 12:00 and 14:00 to 18:00 Beijing time. No effective date has been announced. Those windows correspond to roughly 01:00 to 04:00 and 06:00 to 10:00 UTC.

Cache hits, before anything else. DeepSeek prices a cache hit input token at $0.0028 per million against $0.14 on a cache miss, a ratio of about 50 to 1. Keeping a long system prompt stable so that it caches saves more than switching to a slightly cheaper model, and it is a change you control entirely.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us