
github.com
September 28, 2026
3 min read
46/100
Summary
A distributed pipeline inference engine on multiple ESP32S3 running 1.58-bit (BitNet) Language model. This project runs a sliced 0.5B LLM across a cluster of 7 ESP32s3. One act as master and others are node. The master node runs the tokenizer and embeding and the other attention layer and MLP ran on the nodes. The master and node communicate through high speed SPI Daisy-Chain. ┌───────────────────...