Attention Quantization for Tabular Foundation Models copertina

Attention Quantization for Tabular Foundation Models

Attention Quantization for Tabular Foundation Models

Ascolta gratuitamente

Vedi i dettagli del titolo

Offerte di stagione | 0,99 €/mese per i primi 3 mesi

A seguire 9,99 €/mese – si applicano condizioni. Puoi disdire mensilmente.
Tabular foundation models are served differently from chatbots, so the usual speedups don't carry over. Here the bottleneck is the attention calculation. Converting queries, keys and values to FP8 gave up to 1.7 times the speed with no meaningful accuracy loss on TabPFN-v3 and TabICLv2. The catch: quantization error on test rows has to match the training rows, or accuracy drops sharply. Authors: Jonas M. Kübler, Benjamin Jäger, Klemens Flöge, Noah Hollmann, Frank Hutter Paper: https://arxiv.org/abs/2609.13031v1
adbl_web_anon_alc_button_suppression_t1
Ancora nessuna recensione