HOME/SOLUTIONS/AI Fabric

AI FABRIC

The GPU servers arrived; the switching didn't keep up. Training jobs stall on communication, storage can't feed the cluster, and the next expansion doubles the problem. You need a lossless Ethernet fabric engineered for the cluster you actually run.

WHAT YOU WILL HAVE AT THE END

  • Non-blocking 100/400/800G fabric sized to your GPU count
  • Lossless RoCEv2 transport — PFC/ECN engineered end to end
  • Storage fabric that feeds the cluster at training speed
  • A design that scales to the next GPU tranche without re-architecture
  • Telemetry per port and per queue from day one
  • Documented fabric you can reason about at 2 a.m.
SPINE 1400G · SONiCSPINE 2400G · SONiCLEAF 1AIS/DCS400GGPU RACK8× GPULEAF 2AIS/DCSGPU RACK8× GPU100G RoCEv2 · LOSSLESSLEAF 3AIS/DCSGPU RACK8× GPULEAF 4AIS/DCSSTORAGENVMe PARALLEL FSNON-BLOCKING · POD-REPEATABLE · PFC/ECN ENGINEERED END-TO-ENDREFERENCE DESIGN
Leaf / spineAS7726-32X (100G) · DCS-series 400G platforms400/800G SKUs confirmed per cluster size at design time
NOSSONiC (Enterprise or Community)
Optics / DACET7402-SR4/CWDM4 (100G) · ET7302 (25G server)800G optics specified at design from Edgecore's transceiver line

WHAT WE DELIVER

  1. 01Cluster workload profile and fabric sizing
  2. 02Engineered design + BOM
  3. 03Hardware supply
  4. 04Fabric bring-up incl. RoCE tuning
  5. 05Acceptance: measured all-to-all performance against the design targets
  6. 06Runbook + documentation handover
  7. 07Support tiers

SIZING & TIMELINE

BANDSCALETIMELINE
S≤64 GPUs, 1–4 racks3–5 weeks
M≤512 GPUs6–10 weeks
L512+ GPUs, multi-pod10–16 weeks, phased
  • Outright purchase
  • เช่าใช้ 36-month service agreement
  • Design-only engagement if procurement runs elsewhere

PROVEN IN THE FIELD

Edgecore SONiC fabrics carry production AI training workloads today. Photon State designs and delivers this solution in Thailand on the same platform.

Representative venue image

SAKURA Internet (SAKURAONE)

JAPAN · 2025

SAKURA Internet built its 800-GPU SAKURAONE AI cluster on Edgecore AIS800-64O 800GbE switches running SONiC, arranged as a rail-optimized leaf-spine fabric with RoCEv2 in place of InfiniBand. The system reached #49 on the June 2025 Top500 as the only top-100 machine on a fully open Ethernet networking stack, published by SAKURA's own researchers at MLSys 2026.

Representative venue image

HyperVerge

INDIA · 2026

AI verification company HyperVerge deployed Edgecore high-bandwidth data center switches running SONiC to interconnect its GPU training cluster and a clustered NVMe parallel file system for face-recognition, OCR, and custom LLM workloads. The company reports its 100G RoCEv2 lossless fabric delivers the pooled storage at close to local-NVMe latency, and that open standards let its team self-configure and monitor the network quickly.

EDGECORE DEPLOYMENT · delivered with Broadcom (silicon)SOURCE ↗

QUESTIONS BUYERS ASK

Ethernet or InfiniBand?
We design lossless Ethernet (RoCEv2) on open hardware — the approach Edgecore's public AI deployments run; the trade-offs are stated openly in the design.
Will it work with our NVIDIA/AMD servers?
The fabric is server-vendor-neutral; NIC models and firmware are part of the design inputs.
Who tunes PFC/ECN?
We do, and the acceptance test publishes measured results against agreed targets before handover.
Can we start small?
Yes — the reference design is pod-based; a 2-rack pod carries the same architecture as a 20-rack build.
Lead times on 400G?
Stated per SKU at BOM; the design offers alternates where supply differs.
Support after handover?
Runbook-based support tiers up to managed fabric operation from Bangkok.