Round 9: Tossup 18

The number of these things is the numerator of a ratio that is empirically optimal at about 20, the Chinchilla point. Computations involving the Q, K, and V vectors generated from these things are stored in a KV cache, which improves a metric (-5[1])called the “time to the first” one of these things. The process of converting an input into these things is commonly done using byte-pair encoding. (-5[1])During (10[1]-5[1])the prefill stage, IDs corresponding to these things are embedded into a vector that is then passed through multiple transformer layers. The number of these things that can be processed at once is the context window. The (10[1])cost of (10[1])AI compute is often reported as the price (10[1])per one million of these things. For 10 points, name these units of text that large language models split queries into and try to predict the next one of. ■END■ (10[1])

ANSWER: tokens [accept tokenization or tokenizers; accept time to first token; accept next token prediction or answers referring to trying to predict the next token; prompt on token-to-parameter ratio or tokens-per-parameter ratio; prompt on words or subwords or symbols or characters by asking “those are converted into what things?”]
<GC, Other Science> | Packet K - UCF A, Notre Dame B, Durham A, Michigan State A, Swarthmore A, Cambridge B
= Average correct buzzpoint

Back to tossups