AiToChip
Neural MAX & Neural 1 AI Inference Accelerators
From-scratch design of two custom AI-inference accelerator ASICs (datacenter + edge) with full RTL, simulation, and synthesis flows.
- Architected two custom AI-inference accelerator ASICs from scratch, a 720 mm2 datacenter chip (128 heterogeneous compute slots, 12x HBM3e, 600 W TDP) and a 69.7 mm2 edge chip, modeling 600 W-envelope throughput of 5,500-6,600 tok/s on a 30B-parameter MoE LLM.
- Built a Python RTL generator that compiles a chip specification into 7+ configuration-dependent Verilog modules, enabling one physical die to be re-targeted per model via firmware (SRAM partitioning, precision mode, per-slot power gating).
Verilog / SystemVerilog / Python / Rust / Tcl (synthesis/PnR) / Yosys


