IOSG: From Hot Storage to Cold Memory — Decentralized Storage in the AI-Driven Storage Boom

FIL
AR
Decentralized StorageAI StorageCold DataFilecoinHot DataArweaveHBM
2026-07-29Source: blockweeks.com
IOSG: From Hot Storage to Cold Memory — Decentralized Storage in the AI-Driven Storage Boom

Authors: 0xjacobzhao, IOSG

Recently, "the first domestic storage stock" ChangXin Memory Technologies officially landed on the ChiNext board, and its astonishing performance of a 500% surge ignited the entire market. Although the storage sector is still affected by the aftermath of the correction, AI storage is being heavily revalued by capital amid the current wave of tech narratives. Meanwhile, decentralized storage in the Web3 field has fallen into a long-term silence and disappointment. Why, under the same name of "storage," do market performances differ so drastically? The fundamental answer lies in the complete divergence of the underlying value functions.

The revaluation of storage in the AI era is essentially a carnival about "hot data efficiency," serving the ultimate maximization of "computing utilization" and "commercial monetization." What decentralized storage upholds is the value proposition of "cold data trustworthiness," defending "data fairness, censorship resistance, and the long-term memory of human civilization." The former is an efficiency system for hot data, while the latter is a trust system for cold data. The current capital market undoubtedly stands firmly on the side of "efficiency," but human civilization ultimately still needs an immutable memory foundation. The long-term value of trusted cold storage has never disappeared; it merely lies dormant in the dark side of the cycle, waiting to be repriced by the times.

Why Storage Has Become the Focus of the AI Industry Chain Again

In the traditional IT era, storage was a "capacity business." Enterprise CIOs focused on unit capacity cost, hard drive reliability, disaster recovery plans, archiving strategies, and equipment replacement cycles of 3–5 years. Storage was seen as an accessory that followed server procurement.

The current storage boom is not a traditional cyclical recovery but a repricing of data flow capabilities by AI. In the era of large models, the storage logic has shifted from "capacity first" to "efficiency above all," relentlessly pursuing extreme metrics such as GPU feeding rate, checkpoint writing, and ultra-low latency for RAG. This marks the transformation of storage value from "the final parking lot of data" to "the high-speed channel for data entering computation."

The evolution of resource bottlenecks in AI infrastructure is essentially a battle to fill the "bucket effect." The true utilization of computing power is not a linear superposition of single assets but a strict multiplier effect: real computing utilization = GPU × HBM × DRAM × SSD × Network × File System. A shortcoming in any link can cause the overall computing utilization to collapse. In the AI era, storage has for the first time transformed from a "cost center" into an "efficiency engine." This is the fundamental logic behind the repricing of storage.

HBM

AI Storage Architecture Panorama: From HBM Bandwidth Organ to Data Lake Foundation

AI storage is by no means a simple stack of hardware, but a tightly coupled, hierarchically scheduled complex system. In this system, industrial value and capital focus are highly concentrated on HBM, enterprise SSDs, SSD controllers, NVMe/CXL protocols, and high-performance storage systems. To clearly dissect its value flow, we divide the AI storage architecture into four core layers from top to bottom:

HBM

  • Compute-Near Memory Layer (Bandwidth Core): Dominated by HBM, supplemented by DRAM and CXL memory pooling technology. This layer directly interfaces with GPU/CPU packages or buses, aiming to break the "memory wall" and is the first checkpoint determining whether computing power can be fully unleashed.

  • High-Speed Persistent Storage Layer (IO Hub): The core logic is enterprise SSD = NAND flash + SSD controller + NVMe/PCIe data path. This layer handles high-frequency checkpoint writes, massive training dataset loading, and RAG hot data caching, representing the most significant persistent storage increment in AI data centers.

  • Low-Cost Large-Capacity Storage Layer (Capacity Foundation): Composed of HDDs, cold storage, and data lake archiving systems. Facing exponentially growing multimodal raw data, historical logs, and compliance backups, this layer still provides irreplaceable TCO (Total Cost of Ownership) advantages.

  • AI Storage Systems and Data Software (Scheduling Brain): Including high-performance parallel file systems, distributed object storage, vector databases, and RAG data governance layers. What AI truly consumes is not bare hardware, but data availability efficiently organized, indexed, and permissioned by the software stack.

As an ecosystem extension, decentralized storage does not directly engage in the millisecond-level race of AI hot data. Instead, it anchors on public dataset certification, AI training data provenance, and long-term cold memory archiving, establishing its unique ecological niche as a "trusted cold layer."

HBM: The "Bandwidth Organ" Closest to Computing Power in the AI Storage Chain

High Bandwidth Memory (HBM) is not traditional storage but a high-bandwidth memory layer near the GPU. Its core mission is not to store data but to continuously "feed" data to the computing power at extremely high bandwidth. HBM is the link closest to computing power and the most deterministic in the AI storage chain, directly determining whether the GPU can be "fed" and is the most critical supply chain bottleneck currently.

The core architecture of HBM is "3D DRAM stacking + 2.5D advanced packaging": through TSV vertical stacking and CoWoS heterogeneous integration, it extremely compresses the distance between storage and computation, achieving a generational leap in bandwidth. Its industrial barrier is not just DRAM design, but a system engineering of DRAM process, TSV, ultra-thin stacking, packaging, heat dissipation, testing, and customer certification. Any yield defect in any link can cause the entire HBM stack to be scrapped.

Currently, only the three giants SK Hynix, Samsung, and Micron can stably mass-produce HBM, building a triple moat of top-tier DRAM process, packaging capabilities, and NVIDIA/AMD customer certification.

Generation

Capacity/Die

Bandwidth

Interface Width

Mass Production Time

Key Customers

HBM2e

8–16GB

460 GB/s

1024-bit

2019–2020

Mature generation, used in previous-gen AI accelerators like A100

HBM3

24GB

819 GB/s

1024-bit

2022

SK hynix gained significant first-mover advantage in the H100 cycle

HBM3e

24–36GB

1.2 TB/s

1024-bit

2024–2025

Current main volume driver for AI GPUs. Suppliers are primarily SK hynix, Micron, and Samsung.

HBM4

32–48GB

>2 TB/s

2048-bit

2025–2026

Targeting next-gen AI in 2026,

in mass production introduction/customer qualification phase

HBM4E

64GB+

>2 TB/s

2048-bit+

After 2027

Planning/R&D after 2027

DRAM and CXL: System Memory Foundation and Memory Pooling Engine

HBM addresses the extreme near-GPU bandwidth, DRAM solidifies the server system memory foundation, and CXL attempts to break physical boundaries, reshaping the organization of data center memory resources.

  • DRAM: Mainly carries CPU-side cache, data preprocessing, intermediate state temporary storage, and system operation, serving as the most basic system memory layer of the server. The global DRAM market is highly concentrated among the three giants: SK hynix, Samsung, and Micron; ChangXin Memory Technologies (CXMT) is the core variable for China's DRAM domestic substitution.

  • CXL (Compute Express Link): A new generation cache-coherent interconnect protocol for data centers, aiming to break through the limitations of traditional DIMM slots, local memory capacity, and server memory resource silos, driving the evolution of memory architecture towards expansion, pooling, and sharing. Currently, CXL is still in the early stage from platform support to large-scale deployment, with medium to long-term architectural value; core companies include Astera Labs and Montage Technology.

HBM

Enterprise SSD: Data Hub Built by NAND, Controller, and NVMe

Enterprise SSD is the core high-throughput persistent increment in AI data centers. With extremely high throughput, ultra-low latency, and stable QoS, it continuously "feeds" data to GPUs throughout the entire lifecycle, including training data loading, checkpoint writing, RAG retrieval, inference caching, and log backflow.

In the AI storage architecture, SSD is not an isolated hardware but a highly coupled system, which can be distilled into an industry formula: Enterprise SSD = NAND Flash + SSD Controller + NVMe/PCIe Data Path. The three layers represent independent industry chain segments:

  • NAND Flash (Raw Material Layer): Determines storage density and unit cost. The controller governs performance release and lifespan management. Representative companies: Samsung, SK hynix (Solidigm), Micron, Kioxia, Western Digital, YMTC.

  • SSD Controller (Performance Enabler Layer): Determines performance release, data error correction, QoS stability, and wear leveling. Representative companies: Phison, Silicon Motion, Marvell, Maxio.

  • NVMe/PCIe (Data Path Layer): Determines the transmission efficiency from storage to computation. Combined with GPUDirect Storage technology, it reduces CPU memory bounce buffer and CPU involvement, significantly alleviating I/O bottlenecks. Representative companies: Broadcom, Marvell, Astera Labs.

HDD / Cold Storage / Archive: Low-Cost Foundation of AI Data Lake

AI will not eliminate HDD. With the thirst of multimodal large models for video and image data, as well as the exponential expansion of enterprise compliance logs and historical datasets, the demand for low-cost cold data storage is surging simultaneously. In the AI storage architecture, SSD and HDD collaborate hierarchically based on business value: SSD handles hot data and high throughput, while HDD handles low cost and long-term preservation. Representative companies include Seagate, Western Digital, and Toshiba.

AI Storage Software Stack: Scheduling Hub for Data Availability

What AI truly consumes is never raw disks, but "data services" meticulously organized by the software stack. This architecture transforms underlying hardware into knowledge assets directly callable by upper-layer AI, specifically divided into four layers:

  • High-Performance Storage System (Feeding System): Focuses on concurrent throughput and low latency, solving the "data hunger" problem of GPU clusters through parallel file systems, ensuring rapid flow of training and inference. Representative companies: VAST Data, WEKA, Pure Storage.

  • Object Storage (Raw Data Lake): Centered on Object, Key, and Metadata management, carrying massive unstructured data. It does not pursue extreme low latency but builds a capacity base with low cost and cloud-native characteristics. Representative company: AWS S3.

  • Vector Database (Semantic Index Layer): Vector databases store, index, and retrieve vectors generated by embedding models, enabling AI to accurately locate relevant content from vast knowledge. Representative companies: Pinecone, Milvus.

  • RAG Data Layer (Knowledge Invocation Layer): Goes beyond single retrieval, covering data slicing, cleaning, permission control, and citation tracing, ensuring enterprise data can be safely, accurately, and traceably invoked by large models. Representative company: Databricks.

From AI Hot Storage to Decentralized Cold Memory: Efficiency Maximization vs. Trust Maximization

AI storage is an extreme efficiency-driven system, whose value function focuses on maximizing computational output. HBM bandwidth determines whether GPUs can be fed, SSD throughput determines the read/write efficiency of datasets and checkpoints, and low latency concerns the real-time experience of RAG and inference. These metrics ultimately converge into GPU utilization and cost per Token, directly determining the commercial profitability of AI applications. The ultimate goal of AI storage is not preservation but acceleration, serving productivity.

In contrast, the value function of decentralized storage is entirely different. It asks whether data will still exist in ten years, whether it has been tampered with, and whether it can resist single-point censorship. Through cryptographic proofs and distributed networks, it builds a publicly accessible and permanently preserved public data base. Its ultimate goal is to defend the absolute truth and sovereign independence of data, serving fairness, anti-censorship needs, and civilizational memory.

Dimension

AI Hot Storage

Decentralized Cold Storage

Core Value

Maximize efficiency — accelerate data into computation

Maximize trust — ensure data is not tampered or deleted

Data Type

Hot data / Warm data / Real-time inference

Cold data / Permanent archive / Public memory

Key Metrics

Bandwidth, throughput, latency, GPU utilization

Verifiability, censorship resistance, immutability, permanence

Payment Source

Cloud providers, AI labs, enterprise RAG systems

Public datasets, long-term archives, on-chain applications

Ultimate Goal

Not preservation, but acceleration

Not speed, but trust

Market Status

Hot — Super cycle in progress

Silent — Valuation collapse, narrative drained

AI storage is the "hot storage" that fuels future productivity, while decentralized storage is the "cold memory" that preserves indelible historical records for human civilization. The former serves efficiency, pursuing ultimate speed; the latter serves trust, defending silent memory. The former determines how fast models run, the latter determines whether memories will be deleted. Currently, market mechanisms reward productive efficiency, putting AI storage at the forefront, while decentralized storage seems to be experiencing valuation collapse and narrative draining in silence.

Vision and Reality of Decentralized Storage

There are many decentralized storage projects, but in terms of industry mindshare and ecosystem accumulation, the core representatives have always been Filecoin and Arweave. Although both belong to "decentralized storage," their underlying architectural philosophies are almost two completely different paths — the former uses market-based contracts to approximate AWS elasticity, while the latter uses one-time social contracts to approximate the permanence of a library.

  • Filecoin: Through PoRep and PoSt, it has built the most complete verifiable economic system. It should no longer compete head-on with AWS in consumer-grade cloud storage, but instead pivot to AI data provenance, public dataset hosting, and compliance archiving, providing a verifiable chain for model auditing and copyright proof. The necessary path is to package it as an S3-compatible API and support fiat payment, upgrading from a "cheap storage market" to a "verifiable computing infrastructure."

  • Arweave: With the narrative of "pay once, store forever," it uses Blockweave and SPoRA mechanisms to incentivize miners to preserve and quickly access as much data as possible, especially scarce historical data. Its best position is as the public memory base of humanity — preserving human rights records, war crime evidence, cultural classics, archiving legal and financial history, and providing permanently accessible long-term memory for AI agents. Arweave's value lies not in speed, but in its ability to carry civilizational memory across cycles.

Dimension

Filecoin — Verifiable Storage Market

(Data as of June 2026, source Filfox)

Arweave — Permanent Public Memory Layer

(Data as of June 2026 Source ViewBlock)

Core Philosophy

"Storage is a market" — price discovery, elastic supply, contracts can expire without renewal

"Storage is a public good" — one-time social contract, data exists without relying on any entity's continuous payment

Underlying Structure

Standard blockchain + IPFS content addressing; data and chain stored separately

Blockweave — each new block simultaneously links to the previous block and a random historical "recall block"

Consensus Mechanism

Expected Consensus (EC): Proof of Replication (PoRep) + Proof of Spacetime (PoSt)

SPoRA (Succinct Proof of Random Access, iterated from PoA in 2021)

Storage Proof Logic

PoRep proves the miner generated a unique copy of the data; PoSt continuously proves the copy is still intact

Mining requires proving access to a random recall block — incentivizing miners to retain as much historical data as possible, including unpopular data

Market Structure

Two-layer market architecture: Storage Market (storage deals) + Retrieval Market (retrieval deals), on-chain matching, off-chain data transfer

Single-layer permanent writing; no independent retrieval market, relies on Gateways (e.g., AR.IO) for retrieval services

Payment Model

Storage market / contract system — clients sign periodic lease agreements with storage providers (SPs), pay as needed

Endowment perpetual donation model — one-time payment, funds deposited into an interest pool, theoretically paying miners forever

Data Availability Guarantee

Guarantees integrity (data not tampered), but does not inherently guarantee retrieval speed; additional retrieval services need to be purchased

Guarantees persistence and accessibility; anyone holding the transaction ID can permanently view and download, without relying on the original uploader's wallet

Typical Scenarios

Enterprise cold archiving, compliance evidence storage, verifiable snapshots of AI training data, on-demand elastic storage

Permaweb permanent websites, NFT metadata, historical archives, AI Agent long-term memory (AO computation layer)

Token Model

FIL has a maximum supply cap of 2 billion tokens; actual circulation is affected by block reward release, staking, penalties, and burning mechanisms

AR has a maximum supply of approximately 66 million tokens; the circulating proportion is close to the cap, and additional inflation has a minor impact.

Network Scale

Quality Adjusted Power: approximately 14,848 PiB

Network Size: approximately 20.2 PiB

Cumulative Stored Data

Active deals stored data: approximately 1,110 PiB

Weave Size approximately 0.345 PiB

Miners / Nodes

Approximately 611 active miners

Approximately 100 online nodes

Storage Cost

Filecoin Cloud $2.50/TiB/month/replica

Approximately 10.4–10.7 AR/GiB (≈$20–21)

The dilemma of decentralized storage projects like Filecoin and Arweave is not that their value propositions are wrong, but rather the long-term mismatch between productization, retrieval experience, real demand, and token incentives. This reveals a huge gap from geek ideals to mainstream commercial applications:

  • Mismatch in supply-demand incentives: Early networks represented by Filecoin rapidly expanded capacity through tokens, but failed to build sufficient paid demand, resulting in huge capacity but low utilization and paid conversion. Rewarding "I can store" rather than "I need to store."

  • Lack of enterprise-grade service capabilities: AWS's moat is not hard drives, but the "data operating system" composed of APIs, SLAs, permission management, compliance audits, and technical support. Enterprises buy "peace of mind," not experimental infrastructure that requires handling keys and node selection themselves.

  • Retrieval experience shortcomings: "Storing in" does not equal "stable, low-latency retrieval." Decentralized nodes, complex topology, and lack of unified SLA make it difficult to support AI hot data workflows, more suitable for trusted cold archiving and data provenance.

  • Insufficient privacy compliance: Enterprise private data cannot simply be written into a public permanent network; the right to deletion conflicts with permanent immutability. Decentralized storage is more suitable for public data and long-term archives, not for indiscriminately hosting core private data.

  • Token economy amplifies cycles: Bull market financialization masks insufficient demand; bear market miner ROI decline exposes commercialization shortcomings. Tokens can cold-start supply but cannot automatically create demand and sustainable revenue.

Other decentralized storage projects mostly focus on specific ecosystems or niche tracks: Storj/Sia have weaker cross-cycle industry mindshare and Web3 narrative influence than Filecoin/Arweave; BNB Greenfield/Walrus are tied to specific public chain ecosystems like BNB or SUI; Celestia/EigenDA belong to the data availability (DA) layer, serving Rollup transaction confirmation rather than long-term archiving; 0G and other AI/DA hybrid narrative projects attempt to integrate storage, data availability, computation, and AI agent settlement into an AI-native modular infrastructure, but their real demand, developer adoption, and commercialization closed loop remain to be verified.

Future opportunities for decentralized storage: The long-term pendulum of efficiency and trust

During the technology dividend explosion, capital frantically chases efficiency, assigning high premiums to assets like GPU and HBM, while decentralized storage advocating "trust and fairness" is naturally marginalized. However, the pendulum of history will not stay on the efficiency side forever. Events such as unreasonable bans and content deletions by super platforms, AI copyright lawsuits forcing data provenance, data sovereignty disputes triggered by geopolitical conflicts, data monopolies leading to the disappearance of public archives, and regulatory audit pressure on model training data compliance may all brew a repricing of "trusted storage." The future opportunities for decentralized storage may still demonstrate unique value in the following directions:

  • AI Data Provenance: Combining cryptographic proofs to build "data lineage proofs" to address regulatory and audit pressures.

  • Public datasets and civilization archives: Anchoring censored archives and cultural heritage to build irreplaceable, non-deletable memory.

  • Trusted archiving and compliance evidence: Achieving trusted self-certification through hash-based evidence storage, providing high-level digital notarization.

  • Integration of ZK/TEE/DID technologies: Resolving privacy tensions, upgrading from a single "storage protocol" to a "trusted data infrastructure."

  • Invisible product route: Providing S3-compatible APIs and fiat billing, allowing users to directly purchase "trusted archiving" services.

AI storage and decentralized storage: one pursues extreme efficiency, providing fuel for our journey to the future; the other defends silent memory, preserving our right to look back at the past. The current market unreservedly rewards efficiency, making decentralized storage seem quiet or even collapsing. But when the AI era further amplifies data monopolies, copyright disputes, and the fragility of historical memory, decentralized storage may usher in a value reassessment as a "trusted cold layer." Those memories that cannot be easily erased by platforms, companies, or any single power may evolve from idealistic romance, from fringe belief, into essential infrastructure.