Architect Labs Says Its AI Took Redwood From Chip Spec to FPGA in Two Weeks

The reported speedup could let hardware teams revise designs around changing AI workloads, but Redwood’s performance case remains a projection until it reaches fabricated silicon.

By 2 min read
Architect Labs Says Its AI Took Redwood From Chip Spec to FPGA in Two Weeks
Architect Labs Says Its AI Took Redwood From Chip Spec to FPGA in Two Weeks

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
Architect Labs says it has taken Redwood from a written chip specification to a working FPGA implementation in under two weeks. Two human architects supplied the high-level design; the company says its platform handled performance modeling, RTL generation, testing, verification, firmware, and software kernels without intervention below that level. When the specification changes, Architect Labs says a revised design can return to the FPGA in less than 48 hours. It also reports cutting optimization runs from 15 hours to between 15 and 30 minutes. The current result is a prototype, not a validated chip advantage. Running the Qwen3 0.6B model, Redwood Nano averaged 12.1 tokens per second on an FPGA at 250 megahertz. Nvidia’s Jetson Orin Nano averaged 28 tokens per second in the cited comparison, at 1,020 megahertz. Architect Labs’ advantage appears only in its projection for a future Samsung eight-nanometer-class chip: 49 tokens per second at 1.335 watts, or a modeled 3.4-times improvement in performance per watt. Those figures come from FPGA data, not fabricated silicon. The company says the first RTL design sent from simulation to FPGA had no bugs, with 95 percent code and functional coverage, but automation is not qualification. The important next steps are scaling Redwood to larger models and fabrics, then physical design, a planned TSMC tapeout, and post-silicon validation. The question is whether this faster design loop can survive contact with a real chip.

Story brief

3 key points

Architect Labs has demonstrated a potentially much shorter hardware iteration loop, but not a validated chip breakthrough. Its Redwood Nano FPGA prototype ran Qwen3 0.6B at 12.1 tokens/sec, behind Jetson Orin Nano’s 28; the company models 49 tokens/sec at 1.335W for a future Samsung 8nm-class chip. The platform reportedly took a specification through RTL, verification, firmware, and kernels in under two weeks, with...

  1. 01

    Two human architects supplied the high-level specification; Architect Labs claims no intervention below that level.

  2. 02

    FPGA optimization runs reportedly dropped from 15 hours to 15–30 minutes.

  3. 03

    Architect Labs reports 95% code and functional coverage, with no bugs in the first RTL design sent to FPGA.

Architect Labs says its platform turned a written specification into a working FPGA implementation of its Redwood inference-chip architecture in under two weeks. The result points to a faster way to revise specialized AI hardware, but the company’s headline performance advantage depends on an unbuilt chip.

The faster loop

According to the company’s researchers, two human architects supplied a high-level specification, then the Architect Labs Platform handled performance modeling, hardware-description generation, testing, verification, firmware and kernel generation. Architect Labs says there was no human intervention below that specification.

The more consequential claim is not simply that an AI helped draft hardware, but that it can keep the design moving. After verified specification changes, Architect Labs says a revised design returned to the FPGA in less than 48 hours. The company contrasts that with conventional chip programs, where stages are commonly frozen as work proceeds and later changes may be deferred.

A measured prototype and a projected challenger

Redwood’s reported FPGA result is not yet a speed victory over Nvidia. In the company’s test on Qwen3 0.6B, Redwood Nano averaged 12.1 tokens per second at 250 MHz, versus 28 tokens per second for Nvidia’s Jetson Orin Nano at 1,020 MHz.

The Nvidia advantage reverses only in Architect Labs’ estimate for a Samsung 8nm-class implementation. The company projects 49 tokens per second, roughly half the Jetson baseline’s power draw, and a 3.4-times improvement in performance per watt. Those figures are modeled from FPGA results, not measurements from a fabricated Redwood chip.

Automation is not qualification

Architect Labs says every Redwood block reached 95% code and functional coverage without human verification engineers, and that testing found no bugs in the first RTL design sent from simulation to FPGA. It also says moving optimization runs onto the FPGA reduced their reported runtime from 15 hours to 15–30 minutes.

The decisive test is still ahead: physical design, tapeout and post-silicon validation. Architect Labs says it plans to scale Redwood to larger models and fabrics and is working toward a TSMC tapeout. If the design survives that path, the company argues that hardware teams could update chips around changing workloads instead of locking in assumptions years before deployment.

Sources

  1. venturebeat.comAI shrinks chip design cycle to weeks | VentureBeat

Loading discussion...

YOUR READING SPACE

Notifications