Tokenisation via Convex Relaxations
Jan Tempus ⋅ Philip Whittington ⋅ Craig Schmidt ⋅ Dennis Komm ⋅ Tiago Pimentel
Abstract
Tokenisation is an integral part of the current NLP pipeline. Current tokenisation algorithms such as BPE and Unigram are greedy algorithms---they make locally optimal decisions without considering the resulting vocabulary as a whole. We instead formulate tokeniser construction as a linear program and solve it using convex optimisation tools, yielding a new algorithm we call ConvexTok. We find consistent improvements to intrinsic tokenisation metrics, bits per byte (BpB), and (less consistently) some downstream tasks. Furthermore, ConvexTok allows the user to certify how far their tokeniser is from optimal via a lower bound, and we empirically found it to be within 1\% of optimal at common vocabulary sizes.
Successful Page Load