A running set of technical posts on efficient transformer systems, scientific machine learning, and implementation details.
FLARE++: Stop compressing every field the same way
FLARE made global attention affordable on million-point meshes by routing $N$ tokens through $M \ll N$ learned latent queries. Once trained, though, those $M$ queries are frozen: the same compression template serves every geometry, every boundary condition, and every flow field. FLARE++ keeps FLARE’s low-rank structure and linear cost, and lets the current input shape the queries that compress it. This post walks through the idea. The paper is FLARE++: Low-rank attention with attention-synthesized routing (with Sri Datta Ganesh Bandreddi, Jessica Zhang, and Burak Kara), and the code is in FLARE.py. If you have not read the FLARE post, start there; this one builds directly on it. ...