Retrofitting language models to operate over bytes
MainRecent progress in AI has been driven by end-to-end deep learning systems that learn representations directly from data. Large language models (LLMs) exemplify this trend, achieving strong capabilities by training on massive collections of text3,4. However, despite their apparent generality, contemporary LLMs are not fully end-to-end: before learning can begin, text must first be mapped to a sequence of discrete units called tokens. The choice of tokens, although sometimes overlooked, fundam...
Read more at nature.com