Round 12: Tossup 18

The number of these things is the numerator of a ratio that is empirically optimal at about 20, the Chinchilla point. Computations involving the Q, K, and V vectors generated from these things are stored in a KV cache, which improves a metric called the “time to the first” one of these things. The process of converting an input into these things is commonly done (-5[1])using byte-pair encoding. During the prefill stage, IDs corresponding to these things are embedded into a vector that is then passed through multiple transformer layers. The number of these (10[1])things that can be processed at once is the context window. The cost of AI compute is often reported (-5[1])as the price per one million of these things. (-5[2])For 10 points, name these units of text (10[1])that large language models split queries into and try to predict the next one of. ■END■ (10[2]0[5])

ANSWER: tokens [accept tokenization or tokenizers; accept time to first token; accept next token prediction or answers referring to trying to predict the next token; prompt on token-to-parameter ratio or tokens-per-parameter ratio; prompt on words or subwords or symbols or characters by asking “those are converted into what things?”]
<GC, Other Science> | Packet K - UCF A, Notre Dame B, Durham A, Michigan State A, Swarthmore A, Cambridge B
= Average correct buzzpoint

Back to tossups