Distributed AI Shared Memory and Data Distribution Fabric

A Tiered Memory and Intelligent Data-Distribution Architecture for High-Performance AI Infrastructure

Technology Concept Paper

Author: Christopher Soans
Date: August 2026

Download PDF

Distributed AI Shared Memory and Data Distribution Fabric
Figure 1. Distributed AI Shared Memory and Data Distribution Fabric

Executive Summary

Modern artificial intelligence infrastructure increasingly depends on large clusters of GPUs interconnected by high-speed local and network fabrics. Although GPU computational performance has increased dramatically, overall system performance is determined not only by how quickly GPUs can perform calculations, but also by how efficiently required data can be stored, located, transferred, shared, and reused.

Current GPU servers generally provide each accelerator with dedicated high-bandwidth memory (HBM). Technologies such as NVLink provide high-bandwidth communication between GPUs, while RDMA-capable network fabrics enable efficient communication between servers. These architectures provide substantial performance, but data that must be consumed by multiple GPUs may still need to be copied, transferred, replicated, or independently retrieved. At cluster scale, repeated movement of common data can consume memory capacity, interconnect bandwidth, network bandwidth, and energy while leaving expensive computational resources waiting for information.

This paper proposes a Distributed AI Shared Memory and Data Distribution Fabric (DASM-DDF) that introduces shared memory and intelligent data-distribution tiers at both the server and cluster levels.

Within an AI server, GPUs would retain dedicated local HBM for latency-sensitive working data while gaining access to a high-bandwidth shared-memory tier intended for information used by multiple GPUs.

Across the AI cluster, multiple dedicated Memory Nodes would provide distributed high-speed shared-data storage. These nodes would work with RDMA-capable networking, multicast distribution, local caching, object/version tracking, and delta-based updates to reduce unnecessary data movement.

An AI Intelligence/Control Network could coordinate the architecture by determining what information exists, where it resides, which version is available, which systems require it, and the most efficient method of delivering or reconstructing it.

The objective is not simply to create faster memory. It is to increase the amount of useful AI computation performed per unit of infrastructure, network bandwidth, time, and energy.

1. The Data-Movement Challenge in AI Infrastructure

GPU performance is often measured in computational throughput, but computation represents only one component of an AI workload.

A GPU cannot perform useful calculations unless the required information is available when it is needed.

AI workloads may require movement of:

  • model parameters;
  • tensors;
  • activation data;
  • gradients;
  • embeddings;
  • training samples;
  • inference context;
  • intermediate computational results;
  • reusable mathematical representations;
  • collective-communication data;
  • and other shared computational objects.

When required information is unavailable locally, the GPU may have to wait while it is retrieved from another memory domain or another server.

At scale, even relatively short periods of waiting can become important because the infrastructure supporting an idle or underutilized GPU continues consuming resources.

The GPU itself consumes power, while memory, CPUs, DPUs or SmartNICs, network interfaces, switches, power-conversion systems, and cooling infrastructure continue operating.

The infrastructure problem therefore becomes:

How can required data be placed as close as practical to the computation that needs it while avoiding unnecessary duplication, retransmission, and recomputation?

2. Existing GPU Memory Architecture

A conventional AI server generally contains several distinct memory domains.

The CPU accesses system memory, normally DDR-class RAM, while each GPU typically contains its own local high-bandwidth memory.

A simplified architecture is:

                 CPU
                  |
             System RAM
                  |
              PCIe/CXL
                  |
       -----------------------
       |          |          |
      GPU 0      GPU 1      GPU 2
       |          |          |
      HBM        HBM        HBM
       |          |          |
       ------ NVLink --------

Local HBM provides extremely high bandwidth and is therefore appropriate for information being actively processed by a GPU.

High-speed GPU interconnects allow GPUs to exchange information much more efficiently than conventional host-mediated transfers.

Nevertheless, physical memory remains a finite and valuable resource, and logically shared data may still be physically distributed or replicated.

This creates an opportunity for an additional memory tier.

3. Proposed Intra-Server Shared HBM Tier

The first component of the proposed architecture is a shared high-bandwidth memory tier within the AI server.

Dedicated GPU HBM would not be eliminated.

Instead, the system could provide:

              Shared High-Bandwidth Memory
                         |
                 Memory Fabric
              /          |          \
           GPU 0       GPU 1       GPU 2
             |            |           |
         Local HBM    Local HBM    Local HBM

Local HBM would remain appropriate for:

  • private working data;
  • latency-critical tensors;
  • frequently accessed GPU-specific information;
  • registers and cache backing;
  • and other high-intensity local workloads.

Shared memory would instead be optimized for information consumed by multiple GPUs.

Potential examples include:

  • common model parameters;
  • shared tensors;
  • reference dictionaries;
  • reusable computational results;
  • common inference information;
  • shared lookup structures;
  • and intermediate data required by multiple accelerators.

4. Dedicated Shared-HBM Memory Fabric and Parallel Memory Controllers

The effectiveness of an intra-server shared-HBM tier depends not only on the performance of the HBM itself, but also on the architecture connecting that memory to the GPUs.

A shared-memory implementation could provide little benefit if multiple GPUs must compete across a constrained common bus. Even when the HBM devices themselves provide extremely high aggregate bandwidth, overall performance would be limited by the narrowest component in the data path.

Conceptually:

Effective Shared-Memory Bandwidth =
Minimum of:

• HBM bandwidth
• memory-controller throughput
• shared-memory fabric bandwidth
• GPU memory-interface bandwidth
• contention-limited throughput

For this reason, the proposed architecture should provide the shared-HBM subsystem with a dedicated high-bandwidth memory fabric and multiple parallel memory controllers rather than relying exclusively on the GPU-to-GPU interconnect.

A simplified architecture would be:

                    SHARED HBM POOL

          HBM0      HBM1      HBM2      HBM3
            |         |         |         |
           MC0       MC1       MC2       MC3
            |         |         |         |
        +---+---------+---------+---------+---+
        |       DEDICATED MEMORY FABRIC       |
        +---+---------+---------+---------+---+
            |         |         |         |
          GPU 0     GPU 1     GPU 2     GPU 3

The memory controllers could independently service requests to different HBM banks or partitions, allowing multiple GPUs to retrieve shared information concurrently.

This parallel architecture is important because simply increasing HBM capacity does not guarantee that multiple accelerators can consume the additional memory efficiently.

Separation of Memory and GPU-to-GPU Traffic

The dedicated memory fabric could also prevent shared-memory traffic from unnecessarily competing with GPU-to-GPU communication.

Conceptually, each GPU could have separate high-bandwidth paths:

                           GPU
                     /     \
                    /       \
           GPU Interconnect  Shared-Memory Fabric
            NVLink/etc.             |
                 |                   |
             Other GPUs         Shared HBM

This could allow the GPU to exchange information with other accelerators while simultaneously retrieving shared objects through the dedicated memory subsystem.

Such separation could be especially valuable during communication-intensive AI workloads where both GPU-to-GPU synchronization and shared-data retrieval occur concurrently.

Parallel HBM Organization

The shared-memory pool could itself be divided into multiple HBM partitions, each served by one or more memory controllers.

For example:

                  SHARED HBM POOL

        HBM0 HBM1 HBM2 HBM3 HBM4 HBM5 HBM6 HBM7
          |    |    |    |    |    |    |    |
         MC   MC   MC   MC   MC   MC   MC   MC
          |    |    |    |    |    |    |    |
          +----+----+----+----+----+----+----+
                         |
               Shared-Memory Fabric
                         |
          +------+------+------+------+
          |      |      |      |      |
        GPU0   GPU1   GPU2   GPU3   GPU4 ...

Large shared objects could potentially be striped across multiple HBM partitions so that different portions can be accessed concurrently.

Conceptually:

Shared Object X

Block 0  -> HBM0
Block 1  -> HBM1
Block 2  -> HBM2
Block 3  -> HBM3
Block 4  -> HBM4

...

This could increase aggregate throughput when the workload and memory-access pattern permit parallel retrieval.

The memory-placement system would need to balance striping against locality and contention. Small objects, highly localized objects, and frequently modified information might benefit from different placement strategies.

5. Memory-Level One-to-Many Distribution

A shared-HBM architecture introduces another optimization opportunity when several GPUs request identical information.

A straightforward shared-memory system might independently service each request:

GPU 0 -> Read Object 7421
GPU 1 -> Read Object 7421
GPU 2 -> Read Object 7421
GPU 3 -> Read Object 7421

Even though the object resides in shared memory, this could still cause the HBM subsystem to repeatedly retrieve identical data.

A more efficient shared-memory fabric could recognize simultaneous or closely related requests for the same immutable or versioned object.

The memory subsystem could retrieve the required information once and replicate it within the shared-memory fabric:

GPU0 --+
GPU1 --+
GPU2 --+--> Request Object 7421
GPU3 --+
             |
             v
       Memory Controller
             |
        One HBM Read
             |
             v
      Memory-Fabric Replication
             |
       +-----+-----+-----+
       |     |     |     |
       v     v     v     v
     GPU0  GPU1  GPU2  GPU3

This creates an intra-server one-to-many distribution mechanism conceptually similar to multicast networking.

Rather than repeatedly reading identical data from HBM for every GPU, replication occurs closer to the consumers at the point where their data paths diverge.

This capability would be most practical for immutable or explicitly versioned shared objects, where all receiving GPUs are known to require identical information.

6. Hierarchical One-to-Many Data Distribution

The memory-level replication mechanism can be combined with the network multicast architecture proposed for the Distributed AI Shared Memory and Data Distribution Fabric.

The result is a hierarchical data-distribution model.

At the cluster level:

                    Memory Node
                         |
                    Read Object X
                         |
                  AI Data Fabric
                         |
                  Network Multicast
                         |
          +--------------+--------------+
          |              |              |
          v              v              v
       Server A       Server B       Server C

Within each server:

                   Received Object X
                         |
                  Shared-HBM Fabric
                         |
                Memory-Level Distribution
                         |
              +----------+----------+
              |          |          |
              v          v          v
            GPU 0      GPU 1      GPU 2

The complete data path could therefore become:

                  DISTRIBUTED MEMORY NODE
                           |
                     Object 7421
                           |
                    One Source Read
                           |
                           v
                 NETWORK MULTICAST
                           |
              +------------+------------+
              |                         |
              v                         v
          GPU Server A              GPU Server B
              |                         |
        Shared-HBM Fabric          Shared-HBM Fabric
              |                         |
       One-to-Many Replication    One-to-Many Replication
              |                         |
        +-----+-----+             +-----+-----+
        |     |     |             |     |     |
        v     v     v             v     v     v
       G0    G1    G2            G0    G1    G2

This provides a consistent architectural principle across both the server and the cluster:

Retrieve common data as few times as practical and replicate it as close as possible to the point where the consumers’ paths diverge.

At the cluster level, switches perform efficient network replication.

At the server level, the shared-memory fabric performs efficient accelerator-level replication.

The approach could reduce redundant source reads as well as redundant transmissions.

7. Shared-Memory Fabric Scheduling

Because multiple GPUs may access shared memory simultaneously, the shared-memory subsystem would require intelligent scheduling and arbitration.

Potential considerations include:

  • GPU request priority;
  • memory-controller utilization;
  • HBM-bank availability;
  • object location;
  • object popularity;
  • sequential versus random access;
  • simultaneous requests for identical objects;
  • latency sensitivity;
  • GPU workload phase;
  • available GPU-local HBM;
  • and memory-fabric congestion.

The shared-memory controller could distinguish between requests that should be independently serviced and requests that can be combined.

For example:

GPU 0 requests Object A
GPU 1 requests Object B
GPU 2 requests Object C

-> Parallel memory-controller access

while:

GPU 0 requests Object X
GPU 1 requests Object X
GPU 2 requests Object X

-> Combined read + fabric replication

The objective would be to maximize useful memory bandwidth rather than merely maximize the number of memory transactions.

8. Relationship to GPU-Local HBM

The addition of a shared-HBM subsystem should not eliminate GPU-local HBM.

The two memory types serve complementary purposes.

                    GPU
                  /     \
                 /       \
          Local HBM     Shared HBM
             |              |
        Private/Hot     Common/Shared
           Data            Data

GPU-local HBM remains preferable when information:

  • is accessed extremely frequently;
  • is GPU-specific;
  • requires the lowest practical latency;
  • changes frequently;
  • or would cause excessive shared-fabric traffic.

Shared HBM becomes attractive when:

  • multiple GPUs require identical information;
  • the information is large;
  • maintaining multiple copies wastes capacity;
  • the information is relatively stable or versioned;
  • or repeated GPU-to-GPU transfers would otherwise be required.

The architecture therefore creates a memory hierarchy rather than replacing one memory technology with another.

9. Dynamic Promotion and Replication

Frequently accessed shared information could still be promoted into GPU-local HBM when doing so is more efficient.

For example:

Initial state:

Shared HBM -> Object X
                  |
              GPU accesses



High reuse detected:

Shared HBM -> Object X
                  |
             Local copy
                  |
             GPU Local HBM

Conversely, information that becomes less frequently accessed could be removed from GPU-local memory while remaining available in shared HBM.

The system could therefore dynamically determine whether an object should be:

  • shared only;
  • locally replicated;
  • striped;
  • cached;
  • or removed.

This makes memory placement dependent on actual workload behavior rather than static allocation.

10. Hardware Constraints and Practical Limits

The architecture does not imply unlimited shared-memory bandwidth.

Physical constraints will ultimately limit performance, including:

  • HBM interface bandwidth;
  • memory-controller throughput;
  • semiconductor package area;
  • interposer or advanced-package routing;
  • SerDes capacity where applicable;
  • signal integrity;
  • power delivery;
  • thermal dissipation;
  • arbitration overhead;
  • GPU interface bandwidth;
  • and simultaneous memory-access patterns.

Adding additional memory controllers can increase parallelism only while the rest of the system can support the resulting traffic.

Similarly, increasing memory-fabric bandwidth provides diminishing benefits once HBM, GPU interfaces, or workload characteristics become the limiting factor.

The architecture should therefore be designed as a balanced system, with memory capacity, memory-controller bandwidth, fabric bandwidth, GPU bandwidth, and workload demand engineered together.

11. Updated Architectural Principle

With the dedicated shared-memory fabric, the Distributed AI Shared Memory and Data Distribution Fabric can operate at three major locality levels:

LEVEL 1 — GPU

GPU Local HBM
     |
Lowest-latency working data


LEVEL 2 — AI SERVER

Shared HBM Pool
     |
Multiple Memory Controllers
     |
Dedicated Shared-Memory Fabric
     |
Memory-Level One-to-Many Distribution
     |
Multiple GPUs


LEVEL 3 — AI CLUSTER

Distributed Memory Nodes
     |
AI High-Speed Data Fabric
     |
Network Multicast / RDMA
     |
Multiple AI Servers

The AI Intelligence/Control Network can coordinate all three tiers.

For any required object, the infrastructure could determine:

  1. Is the object already in GPU-local HBM?
  2. Is it available in server shared HBM?
  3. Is another GPU-local copy more efficient to access?
  4. Is it available from a nearby Memory Node?
  5. Do multiple GPUs need the same object?
  6. Do multiple servers need the same object?
  7. Can the object be distributed using memory-level replication?
  8. Can it be distributed using network multicast?
  9. Can only a delta or residual be transferred?
  10. Would local recomputation be cheaper than retrieval?

The resulting objective is not simply maximum memory bandwidth.

It is:

Minimum data movement required to keep the maximum amount of GPU computation productive.

This refinement strengthens the overall Distributed AI Shared Memory and Data Distribution Fabric by ensuring that the shared-memory layer itself is designed to scale with the accelerators consuming it.

A dedicated, highly parallel shared-memory fabric prevents shared HBM from merely becoming another centralized resource through which all GPUs must compete. Multiple memory controllers provide parallel access, while hardware-assisted one-to-many distribution allows common data to be read once and efficiently delivered to multiple accelerators.

Combined with distributed Memory Nodes and network-level multicast, the architecture establishes the same principle from memory chip to AI cluster:

Store information intelligently, retrieve it from the nearest efficient source, and replicate common information only at the point where consumer paths diverge.

12. Shared References Instead of Repeated Copies

Consider an object required by four GPUs.

A conventional implementation might result in multiple physical copies:

GPU 0 HBM: Object X
GPU 1 HBM: Object X
GPU 2 HBM: Object X
GPU 3 HBM: Object X

A shared-memory implementation could instead maintain:

Shared HBM

Object ID: 7421
Address: Shared Memory Location
Version: 17

        /      |      |      \
     GPU 0   GPU 1   GPU 2   GPU 3

Each GPU could reference the same logical object.

The address or object identifier itself does not eliminate physical data transfer: the actual bytes still have to cross the memory fabric when accessed. The advantage is that unnecessary copies and redundant storage may be reduced.

This distinction is important.

The architecture does not assume that memory references make data movement free. Instead, it attempts to ensure that data is moved only when and where necessary.

13. Memory as a Hierarchical Resource

The architecture can be viewed as a hierarchy:

GPU Registers / Cache
        |
   GPU Local HBM
        |
 Server Shared HBM
        |
 CPU/System Memory
        |
 Cluster Memory Nodes
        |
 Persistent Storage

Different information can be placed according to:

  • access frequency;
  • latency sensitivity;
  • object size;
  • number of consumers;
  • available capacity;
  • network conditions;
  • cost of recomputation;
  • and energy cost.

Very frequently accessed information remains close to the GPU.

Information shared by several GPUs can reside in server-level shared memory.

Information shared across servers can reside in the distributed Memory Node layer.

Cold information can remain in lower-cost storage.

14. Extending Shared Memory Across the AI Cluster

The same principle can be extended beyond an individual server.

Instead of requiring every AI server to independently retrieve or maintain identical shared information, the cluster could contain dedicated Memory Nodes.

              Distributed Memory Nodes

        +-----------+   +-----------+
        | Memory    |   | Memory    |
        | Node A    |   | Node B    |
        | HBM/Fast  |   | HBM/Fast  |
        +-----+-----+   +-----+-----+
              \             /
               \           /
              AI Data Fabric
                    |
          -----------------------
          |          |          |
       Server A   Server B   Server C
          |          |          |
        GPUs       GPUs       GPUs

The Memory Nodes could contain HBM or another suitable class of high-performance memory.

They would function as a distributed shared-data tier rather than as conventional storage servers.

15. Why Multiple Memory Nodes Are Important

A single central memory server would introduce several risks.

It could become:

  • a bandwidth bottleneck;
  • a latency concentration point;
  • a failure domain;
  • a congestion point;
  • or a scaling limitation.

The proposed architecture therefore uses multiple Memory Nodes.

Objects could be:

  • distributed;
  • sharded;
  • replicated;
  • cached;
  • geographically or topologically localized;
  • or dynamically relocated.

For example:

Memory Node 1
Objects 1-5000

Memory Node 2
Objects 5001-10000

Memory Node 3
Replicas of high-demand objects

Memory Node 4
Recently generated shared computational results

Actual placement would be dynamic rather than necessarily based on fixed object ranges.

16. Intelligent Object Directory

A control system would maintain knowledge of distributed objects.

For example:

Object 7421
Version 17

Memory Node A: Present
Memory Node B: Present

GPU Server 11: Present / Version 17
GPU Server 12: Present / Version 17
GPU Server 13: Present / Version 16
GPU Server 14: Absent
GPU Server 15: Present / Version 17

This directory would allow the infrastructure to avoid transmitting information unnecessarily.

Instead of asking only:

Where is Object 7421?

the system could ask:

Which systems already possess Object 7421, which version do they possess, which systems need it, and what is the least expensive way to satisfy the remaining demand?

17. Multicast Distribution of Shared AI Data

When identical information must be delivered to many servers, conventional point-to-point transmission can result in repeated movement of the same data.

Conceptually:

Unicast

Source ---- Object X ---- Server A
Source ---- Object X ---- Server B
Source ---- Object X ---- Server C
Source ---- Object X ---- Server D

A multicast-capable AI fabric could instead allow:

                    +---- Server A
                    |
Memory Node ---- Switch ---- Server B
                    |
                    +---- Server C
                    |
                    +---- Server D

The Memory Node transmits a shared stream, while the network fabric replicates the information where required.

This can reduce repeated transmission across common portions of the network.

Potential benefits include lower:

  • source-memory read demand;
  • source-NIC utilization;
  • serialization overhead;
  • redundant link utilization;
  • and repeated data movement.

18. Multicast as a Selective Tool

The architecture should not assume that multicast is appropriate for every AI communication.

Point-to-point RDMA or other unicast mechanisms remain appropriate when:

  • only one destination requires the data;
  • different destinations require different data;
  • strict connection-level delivery semantics are required;
  • or latency characteristics favor direct communication.

Multicast becomes particularly attractive when many receivers require identical information.

The infrastructure therefore selects between transmission mechanisms rather than replacing unicast universally.

19. Reliable Multicast with Selective Repair

Large-scale AI computation cannot simply assume that every multicast receiver successfully receives every data block.

A practical architecture therefore requires recovery.

One possibility is:

Memory Node
     |
Multicast Object 7421
     |
 -----------------------------
 |        |        |         |
 A        B        C         D
 OK       OK      Missing    OK
                    |
                   NACK
                    |
             Unicast Repair

Rather than every successful receiver generating acknowledgements, a receiver detecting missing or corrupted information could request retransmission.

The system could then use direct unicast or RDMA to repair only the missing data.

This produces a hybrid model:

Multicast for efficient bulk distribution.

Unicast/RDMA for targeted recovery and exceptions.

Forward-error-correction techniques could also be evaluated where appropriate.

20. Direct Placement into Accelerator Memory

Another objective should be minimizing unnecessary CPU involvement.

Where supported by the hardware and software stack, network interfaces, DPUs, or SmartNICs could facilitate direct movement between the network and accelerator-accessible memory.

Conceptually:

Memory Node
     |
High-Speed Fabric
     |
    DPU
     |
Shared/Local GPU Memory
     |
    GPU

The objective is to avoid unnecessary paths such as:

Network
   |
 CPU RAM
   |
 CPU Copy
   |
 GPU Memory

whenever safe direct placement is possible.

This builds upon the general principles demonstrated by modern RDMA and direct accelerator-memory technologies.

21. Delta-Based Distribution

Multicast reduces duplication in transmission, but further efficiencies become possible if receiving servers already contain an earlier version of an object.

Suppose several servers contain:

Object 7421
Version 16

and the Memory Node contains:

Object 7421
Version 17

Instead of redistributing the complete object, the system could distribute only the changes.

For example:

Object: 7421
Base Version: 16
Target Version: 17

Block 1092: Replace
Block 4105: Replace
Block 7221: Apply Delta

The receiving system reconstructs Version 17 locally.

This could substantially reduce network traffic when changes are small relative to total object size.

22. Combining Multicast and Delta Distribution

Delta distribution becomes particularly powerful when combined with multicast.

If 100 servers contain Version 16 and require Version 17, the system could multicast the update once rather than distributing the entire Version 17 object independently 100 times.

               Memory Node
                   |
          Delta v16 -> v17
                   |
               Multicast
                   |
       -------------------------
       |       |       |       |
      S1      S2      S3     ... S100
       |       |       |          |
      v17     v17     v17        v17

This changes the optimization problem from:

How quickly can we transmit the complete object?

to:

How little information must be transmitted for every destination to obtain the required state?

23. Distributed Computation Reference Library

The Memory Node architecture can also support reusable computational references.

Frequently recurring calculations, tensor structures, transformations, or other reusable representations could be stored as reference objects.

This resembles the historical concept of a logarithm book.

Before electronic calculators, complex mathematical operations could sometimes be simplified by looking up previously calculated values rather than recomputing everything manually.

An AI infrastructure could apply a modern computational version of that principle.

Instead of repeatedly transmitting or calculating a large representation:

Large Computation
       |
       |
       v
Large Result

the system could determine whether a known representation already exists:

Reference Object 7421
       +
Small Residual / Delta
       |
       v
Required Result

The value of such an approach would depend heavily on workload regularity, lookup overhead, numerical requirements, and the cost of reconstructing results, so it should be applied selectively rather than assumed to benefit all computation.

24. The AI Intelligence/Control Network

A dedicated AI Intelligence/Control Network could coordinate the shared-memory and data-distribution architecture without carrying the bulk payload itself.

The control layer could maintain information concerning:

  • object location;
  • object version;
  • memory availability;
  • GPU demand;
  • network congestion;
  • multicast membership;
  • server topology;
  • link utilization;
  • memory-node health;
  • cache state;
  • data popularity;
  • expected reuse;
  • and energy or performance characteristics.

The high-speed AI data fabric would carry the actual tensors and other large objects.

This creates separation between:

AI Intelligence/Control Network
        |
Metadata / decisions / coordination
        |
        +----------------------------+
                                     |
AI High-Speed Data Fabric            |
        |                            |
HBM <-> GPU <-> DPU <-> Network <-> Memory Nodes

25. Dynamic Transfer Decisions

For every required object, the system could evaluate several possible actions.

For example:

  1. Already in local GPU HBM?
    |
    YES --> Use locally
            |
           NO
            v
  2. Available in server shared memory?
            |
           YES --> Retrieve locally
            |
           NO
            v
  3. Available on another local GPU?
            |
           YES --> Use GPU fabric
            |
           NO
            v
  4. Available from nearby Memory Node?
            |
           YES --> RDMA/direct transfer
            |
           NO
            v
  5. Needed by many servers?
            |
           YES --> Multicast
            |
           NO
            v
  6. Transfer by unicast/RDMA

Another decision could precede all of these:

Is retrieving the information actually cheaper than recomputing it?

In some situations, recomputation could be faster or more energy efficient than retrieving a distant object.

26. Cost-Aware Data Movement

The scheduler could therefore evaluate a conceptual cost function:

Total Cost =
    Retrieval Latency
  + Network Cost
  + Memory Cost
  + Synchronization Cost
  + Reconstruction Cost
  + Energy Cost

Possible choices could include:

  • calculate locally;
  • retrieve from local HBM;
  • retrieve from shared HBM;
  • retrieve through the local GPU fabric;
  • retrieve from a Memory Node;
  • retrieve from another server;
  • multicast a complete object;
  • multicast a delta;
  • or reconstruct from a reference object.

The optimal choice can change dynamically as workloads and network conditions change.

27. Data Locality and Predictive Placement

The system does not necessarily have to wait for a GPU to request information.

The AI Intelligence/Control system could analyze job execution and anticipate future demand.

If a scheduled computation will shortly require Object 7421 on Servers 10 through 20, the infrastructure could pre-position the object before those GPUs reach that phase.

Current computation
        |
        |     Background prefetch
        | ------------------------>
        |                  Object ready
        v                       |
Next computation --------------> GPU

This could hide some data-transfer latency behind ongoing computation.

Predictive placement would need safeguards so inaccurate predictions do not consume excessive bandwidth or displace more valuable data.

28. Memory-Node Locality

Not every Memory Node should be treated as equally distant.

A topology-aware system could select the closest appropriate source.

For example:

Rack Group A
  |
Memory Node A
  |
GPU Servers

Rack Group B
  |
Memory Node B
  |
GPU Servers

          |
   Higher-Level Fabric
          |
Memory Node C
Shared/replicated objects

Frequently used information could be replicated near consumers.

This creates a distributed hierarchy similar to caching systems but optimized for AI-scale data objects and accelerator access.

29. Preventing Memory Nodes from Becoming Bottlenecks

Memory Nodes would require careful engineering because accelerating the consumers merely moves the bottleneck if the shared-memory layer cannot supply sufficient bandwidth.

Possible approaches include:

  • multiple memory controllers;
  • multiple high-speed NICs;
  • DPU-assisted networking;
  • object sharding;
  • replication of high-demand objects;
  • load-aware routing;
  • topology-aware placement;
  • hot-object caching;
  • multicast offload;
  • and dynamic load balancing.

The system should monitor memory-node utilization just as carefully as GPU utilization.

30. Consistency and Synchronization

Shared memory introduces consistency challenges.

The system must distinguish between objects that are:

  • immutable;
  • append-only;
  • versioned;
  • mutable;
  • synchronized;
  • or GPU-private.

Immutable and versioned objects are particularly attractive because they reduce coherence complexity.

Instead of allowing arbitrary modification of Object 7421, the system might create:

Object 7421 / Version 16
Object 7421 / Version 17

Consumers explicitly identify the required version.

This reduces ambiguity and allows older versions to remain available while active workloads finish using them.

31. Security and Isolation

A shared accelerator-memory infrastructure must not weaken workload isolation.

Memory objects should be associated with:

  • workload identity;
  • tenant identity;
  • access policy;
  • encryption context where required;
  • integrity information;
  • and lifecycle state.

A GPU or server should receive access only to objects authorized for its workload.

DPUs or SmartNICs could potentially enforce access close to the data-transfer path.

The control system should therefore coordinate both where information resides and who is permitted to retrieve it.

32. Fault Tolerance

Memory Nodes should not represent persistent single points of failure.

Important objects could be replicated across multiple nodes:

Object 7421 / v17

Memory Node A: Primary
Memory Node B: Replica
Memory Node C: Replica

If Node A fails, clients could retrieve the object from another node.

For critical active workloads, placement policies could ensure replicas exist across separate failure domains.

33. Potential Energy Benefits

The architecture’s energy benefit does not come merely from transferring information faster.

The larger opportunity comes from reducing:

  • redundant computation;
  • duplicate memory storage;
  • repeated network transmission;
  • unnecessary CPU copying;
  • GPU waiting time;
  • and excessive server requirements.

The relevant measurement should therefore be energy per completed unit of useful AI work, not merely instantaneous power consumption.

If better memory utilization and data distribution allow a workload to complete faster—or allow the same performance objective to be achieved with fewer GPUs—the total infrastructure energy requirement may decline.

Actual savings would need to be measured workload by workload.

34. Potential Reduction in Required Compute Infrastructure

Faster data delivery cannot reduce the fundamental arithmetic requirements of a workload.

However, communication and memory inefficiencies can cause additional GPUs to be deployed merely to achieve a target completion time.

If GPUs spend substantial time waiting for:

  • synchronization;
  • remote data;
  • memory copies;
  • network transfers;
  • or repeated shared objects,

then improving those areas can increase effective GPU utilization.

In communication-bound or memory-bound workloads, higher utilization could potentially allow a given performance objective to be achieved using fewer compute resources.

This would need to be demonstrated through benchmarking rather than assumed universally.

35. Compatibility with Existing Infrastructure

An important design principle should be incremental deployment.

The architecture should build where possible upon existing technologies such as:

  • Ethernet;
  • InfiniBand;
  • RDMA;
  • RoCE;
  • PCIe;
  • CXL where applicable;
  • GPU direct-memory technologies;
  • DPUs and SmartNICs;
  • multicast-capable switching;
  • IGMP/MLD mechanisms where applicable;
  • and existing AI workload orchestration systems.

A first implementation could potentially be software-defined and use existing server memory and high-speed network infrastructure to test the scheduling and distribution concepts before specialized shared-HBM hardware is developed.

36. Possible Deployment Evolution

A practical development path could proceed in stages.

Phase 1 — Software Demonstration

Use existing GPU servers and high-speed networking to test:

  • object identification;
  • version tracking;
  • data reuse;
  • intelligent source selection;
  • multicast distribution;
  • selective retransmission;
  • and delta transfer.

Phase 2 — Dedicated Memory Nodes

Deploy specialized servers containing large quantities of fast memory and high-bandwidth network interfaces.

Phase 3 — Intra-Server Shared Accelerator Memory

Introduce shared high-bandwidth memory accessible by multiple GPUs alongside dedicated GPU HBM.

Phase 4 — Hardware-Assisted Distribution

Move object lookup, multicast handling, integrity verification, and direct placement into DPUs, SmartNICs, switches, or specialized memory controllers where beneficial.

Phase 5 — Fully Integrated AI Data Fabric

Coordinate computation scheduling, memory placement, data distribution, networking, and energy optimization through the AI Intelligence/Control infrastructure.

37. Example End-to-End Operation

Consider a training operation requiring the same large object on 64 GPU servers.

The system determines:

64 servers require Object X / Version 22

40 already have Version 22
16 have Version 21
8 do not have Object X

Instead of transmitting the entire object 64 times:

Step 1: No transmission is required for the 40 current servers.

Step 2: The system determines whether the 16 Version-21 servers can efficiently reconstruct Version 22.

Step 3: A Version-21-to-Version-22 delta is multicast to those 16 servers.

Step 4: The complete Version 22 object is multicast to the eight servers without a valid base version.

Step 5: Receivers verify object integrity.

Step 6: Any missing blocks are repaired through targeted unicast/RDMA.

Step 7: DPUs or SmartNICs place received data into the appropriate accelerator-accessible memory where supported.

Step 8: The directory records which servers now contain Version 22.

Future requests can therefore avoid unnecessary network transfers.

38. Relationship to Collective AI Communication

This architecture should complement rather than attempt to replace optimized collective-communication libraries and mechanisms.

Operations such as AllReduce, AllGather, ReduceScatter, and All-to-All have specific communication requirements.

The proposed shared-data fabric is particularly relevant where:

  • identical objects are repeatedly distributed;
  • objects persist beyond a single collective operation;
  • common data is reused across jobs or computation phases;
  • large immutable objects can be cached;
  • or delta/reference representations significantly reduce required movement.

Benchmarking would determine where shared-memory distribution is superior to existing collective mechanisms.

39. Measuring Success

A prototype should measure more than raw network throughput.

Useful metrics include:

  • GPU compute utilization;
  • GPU stall time;
  • HBM utilization;
  • duplicate memory consumption;
  • bytes transferred per completed job;
  • network-link utilization;
  • Memory Node utilization;
  • multicast efficiency;
  • delta hit rate;
  • object-cache hit rate;
  • average retrieval latency;
  • job completion time;
  • GPUs required for a performance target;
  • server power;
  • network power;
  • and total energy per completed workload.

One particularly useful metric would be:

Useful AI computation per joule of total infrastructure energy.

40. Key Engineering Challenges

The architecture introduces significant engineering questions that require experimental validation.

These include:

  • providing sufficient shared-memory bandwidth;
  • avoiding contention;
  • maintaining low latency;
  • object consistency;
  • synchronization;
  • cache invalidation;
  • reliable multicast;
  • congestion control;
  • access security;
  • hardware interoperability;
  • workload scheduling;
  • memory fragmentation;
  • topology awareness;
  • failure recovery;
  • and determining when data retrieval is actually preferable to recomputation.

These challenges argue for incremental experimentation rather than immediate replacement of established GPU architectures.

41. Broader Architectural Direction

Today’s AI clusters can be viewed largely as collections of powerful GPUs connected by increasingly fast networks.

A future architecture could instead treat the entire cluster as a coordinated computational and memory system.

                    AI Intelligence
                    / Control Plane
                           |
          -----------------------------------
          |                                 |
     Compute Placement                 Data Placement
          |                                 |
          v                                 v

   +-------------+                  +---------------+
   | GPU Compute | <--------------> | Shared Memory |
   | Resources   |    Data Fabric   | Resources     |
   +-------------+                  +---------------+
          |                                 |
          +-------------+-------------------+
                        |
                 Network Fabric

The scheduler would no longer ask only:

Which GPU should perform this calculation?

It would also ask:

Where is the required information, where should it be placed, and what is the least expensive way to make it available?

That represents a transition from compute-centric AI infrastructure toward compute-and-data-aware AI infrastructure.

Conclusion

The continued scaling of AI infrastructure requires more than increasingly powerful GPUs. It requires increasingly efficient coordination between computation, memory, and networking.

The proposed Distributed AI Shared Memory and Data Distribution Fabric introduces a tiered architecture combining:

  • dedicated GPU HBM;
  • intra-server shared high-bandwidth memory;
  • distributed Memory Nodes;
  • intelligent object and version tracking;
  • RDMA-based point-to-point communication;
  • multicast distribution;
  • selective retransmission;
  • delta-based updates;
  • reusable computational references;
  • topology-aware caching;
  • predictive data placement;
  • and AI-assisted orchestration.

The central principle is straightforward:

Do not move, duplicate, or recompute information unless doing so is more efficient than reusing what already exists.

A future AI infrastructure should know what information it possesses, where that information resides, which version is available, which compute resources will need it, and how it can be delivered with the least overall cost.

By treating memory and data movement as coordinated infrastructure resources alongside GPU computation, this architecture could improve accelerator utilization, reduce redundant network traffic, improve memory efficiency, shorten workload completion times, and potentially reduce the number of compute resources and total energy required for communication- and memory-constrained AI workloads.

The result would be an AI cluster designed not merely to compute faster, but to make more intelligent use of every calculation, every byte of memory, and every byte transferred across the network.

© 2026 Christopher Soans. All rights reserved.

This work is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0).