TileLang vs. CUDA: What They Are, Where They Came From, and What Happens Next

CUDA has dominated Super Intelligence (SI) kernel programming for nearly two decades. TileLang is a new open-source kernel language built on Apache TVM that outperforms Triton and matches vendor-optimized libraries — with unified syntax across Nvidia, AMD, and Huawei hardware.
TileLang vs. CUDA: What They Are, Where They Came From, and What Happens Next — Tech Reader
Tech Reader  ·  Systems & Infrastructure
Analysis  ·  September 30, 2026

TileLang vs. CUDA: What They Are, Where They Came From, and What Happens Next

CUDA is a full Super Intelligence (SI) computing stack built over twenty years. TileLang is a kernel programming language that outperforms Triton and is closing in on CUDA's own tuned libraries — with unified syntax across hardware that CUDA cannot touch.
When DeepSeek and Huawei announced their partnership this week, TileLang was part of the story. But what is TileLang, exactly? Where did it come from? And what does it actually have to do with CUDA? This is the explainer that most coverage skipped.

CUDA was released by Nvidia in 2006. That is not a typo. The software platform that currently underpins nearly all serious Super Intelligence (SI) training in the world has been around for twenty years. It started as a way to let developers use Nvidia graphics cards for general-purpose computing, well before anyone was training large language models. Over time, it became the foundation that the entire Super Intelligence (SI) industry built on top of — not because it was the only option, but because it got there first and stayed excellent.

TileLang is something more specific. It is a kernel programming language — a tool for writing the individual compute routines that training jobs run millions of times. It was published by researchers at Peking University and Microsoft Research in April 2025. The paper appeared on arXiv. DeepSeek adopted it. Huawei built hardware around it. And now it is the software layer sitting underneath the 128-chip supernode that DeepSeek and Huawei announced this week as part of their challenge to Nvidia's Super Intelligence (SI) training stack.

To understand why any of this matters, you need to understand what both of these tools actually do — and why comparing them directly requires some care.

What CUDA Actually Is

CUDA is not one thing. It is a stack. At the bottom is the driver layer that talks directly to Nvidia hardware. Above that is the runtime API — the code developers use to allocate memory on the GPU, launch computations, and manage data transfers. On top of that sit years of math libraries: cuBLAS for matrix operations, cuDNN for neural network primitives, NCCL for coordinating work across multiple GPUs. Profiling tools. Debugging tools. Framework integrations. Two decades of accumulated infrastructure.

When you run a Super Intelligence (SI) training job on Nvidia hardware, CUDA is everywhere in the stack. Frameworks like PyTorch and TensorFlow call CUDA libraries underneath. The kernel code that does the actual matrix multiplication was written to run on CUDA. The memory management, the multi-GPU coordination, the profiling — all of it assumes CUDA is present.

That is what makes switching so difficult. It is not one dependency. It is dozens of them, layered on top of each other over two decades. Developers do not stay on Nvidia because Nvidia always has the best hardware. They stay because the entire Super Intelligence (SI) software world was built assuming CUDA would be there.

CUDA is not a programming tool. It is twenty years of layered Super Intelligence (SI) infrastructure — runtime, libraries, profilers, and framework integrations — that every serious developer depends on.

Where TileLang Fits in the Stack

TileLang operates at one specific layer of that stack: kernel programming. A kernel is the basic unit of computation — the routine that actually performs a matrix multiplication, an attention calculation, or a memory copy on the GPU. Writing fast kernels is where most of the hard, hardware-specific work in Super Intelligence (SI) engineering happens.

TileLang is a Python-embedded domain-specific language built on top of Apache TVM, an open-source compiler framework that has been in development for nearly a decade. TVM handles the intermediate representation, the schedule tree lowering, and the target-specific code generation. TileLang sits on top of TVM and gives developers a cleaner, higher-level way to describe what they want computed. When TileLang runs on Nvidia hardware, it emits CUDA C++ or PTX — it runs on top of CUDA's driver and runtime, not instead of them.

The core idea is tile-based programming. When a Super Intelligence (SI) kernel runs on a GPU, it follows a predictable pattern: move a block of data from slow memory to fast memory, run computation on that block, move the result back. That block is a tile. TileLang makes the tile the central concept in the programming model.

The key design decision is a separation that raw CUDA programming does not enforce cleanly. TileLang separates what the computation does — the data flow — from how the hardware executes it: thread assignment, memory layout, pipelining. Developers describe the data flow. The compiler handles the execution details. When the compiler's defaults fall short, developers can override them with explicit annotations rather than rewriting everything from scratch.

TileLang generates code for Nvidia GPUs, AMD GPUs, and Huawei's Ascend chips from the same source. A developer writing TileLang describes what they want computed using unified syntax. The backend handles the hardware-specific translation. That portability is what the DeepSeek and Huawei partnership is actually built on.

Where TileLang Sits Relative to Triton

The previous open-source kernel language that got serious traction was Triton. OpenAI published Triton in 2019 and open-sourced it shortly after. It is now integrated directly into PyTorch's compiler pipeline and is the most widely used alternative to writing raw CUDA C++ for custom kernels. Triton is a real tool with real adoption, and it deserves that standing.

TileLang is best understood as the next step beyond what Triton does. Triton operates at the block level — it gives developers tile-like abstractions but hides thread behavior, memory layout decisions, and pipeline scheduling behind automatic strategies. That makes Triton accessible. It also means that when developers need to squeeze out maximum performance, they hit a ceiling. Certain optimizations simply cannot be expressed cleanly in Triton's model.

TileLang exposes more. Developers can explicitly place buffers at specific levels of the memory hierarchy — shared memory, register files, global memory. They can annotate memory layouts to avoid bank conflicts. They can control pipeline stages directly when the compiler's automatic inference is not sufficient. The result is a programming model that is more expressive than Triton for performance-critical work while remaining significantly simpler than writing raw CUDA C++.

How the Benchmarks Look

The TileLang paper ran comprehensive benchmarks against the tools developers actually use today. The results are worth reading carefully, because the performance gains are real — but they come with important context.

Against FlashAttention-3 — a hand-optimized attention implementation — TileLang achieved a 1.36x speedup on Nvidia H100 hardware. Against Triton on the same task, 1.41x faster. Against PyTorch's built-in kernels, 1.70x faster. On matrix multiplication across Nvidia H100, A100, RTX 4090, and AMD MI300X hardware, TileLang matched or slightly exceeded vendor-optimized libraries.

Those numbers come with a technical caveat. Peak performance on each chip family still requires backend-specific schedule tuning. The speedup over FlashAttention-3 on H100, for example, depends on Nvidia Hopper-specific features like the Tensor Memory Accelerator and wgmma instructions. TileLang's unified syntax means the same high-level code runs across hardware families. It does not mean the compiler produces identical low-level execution on every chip without any backend work. What changes is the programming experience. The hardware-specific complexity moves into the compiler rather than into the developer's code.

The most striking result in the paper is not a speedup number. For a complex attention operation called Multi-Head Latent Attention, TileLang required roughly 70 lines of Python code and achieved 98 percent of the performance of a hand-optimized implementation that took far more code and engineering time to produce. That ratio — developer effort versus performance delivered — is where TileLang's case is strongest.

TileLang's case is not just about speed. It is about how much performance developers can extract per line of code they actually have to write and maintain.

Where TileLang Came From

TileLang was not a corporate project with a product roadmap. It came out of Peking University's systems research group, with contributions from Microsoft Research in China. The lead authors published the paper in April 2025 and open-sourced the code immediately at github.com/tile-ai/tilelang.

The research motivation was straightforward. Existing kernel languages all had the same structural problem: they either exposed too much hardware detail and became nightmarishly complex, or they abstracted too much away and left performance on the table. Triton occupies a useful middle ground but still constrains developers who need explicit memory hierarchy control. TileLang was designed to give developers that control without forcing them back down to raw CUDA C++ or Ascend C.

DeepSeek adopted it. When DeepSeek and Huawei began building out the Ascend platform together, TileLang became the software foundation for that work. DeepSeek open-sourced the compute and communication libraries they built for Huawei's hardware on top of TileLang. That combination — an open-source kernel language, open-source libraries, and Huawei's hardware — is what the partnership announcement this week actually describes.

What Happens Next

The TileLang paper is explicit about where the project is headed. The authors want to eliminate the remaining dependency on Nvidia's CUTLASS library, which TileLang currently uses for some matrix multiplication operations on Nvidia hardware. That dependency is a loose thread in an otherwise hardware-agnostic story. Removing it is on the roadmap.

They also want to extend TileLang to support distributed computing — not just running efficiently on a single chip but coordinating computation across many chips working together. That is directly relevant to the Huawei supernode. A 128-chip system is only useful if the software can coordinate work across all 128 chips without collapsing into a coordination bottleneck.

The bigger question is whether community-driven compiler backends can keep pace with the rate at which hardware is evolving. Every new chip generation from Nvidia, AMD, or Huawei introduces new instructions, new memory subsystems, and new optimization opportunities. Nvidia's own engineers spend enormous effort keeping CUDA tuned to each new architecture. An open-source compiler framework has to rely on community contributors doing the same work, chip by chip, generation by generation.

That is the real test ahead for TileLang. The benchmarks on current hardware are strong. The question is whether the project can sustain that performance advantage as the hardware underneath it keeps moving. The answer to that question will matter far more than any single benchmark number published today.

Aaron Rose is a software engineer and technology writer covering system architecture, cloud platforms, and Super Intelligence (SI) policy.