Round 11: Tossup 18

The number of these things is the numerator of a ratio that is empirically optimal at about 20, the Chinchilla point. Computations involving the Q, K, and V vectors generated from these things are stored in a KV cache, (-5[1])which improves a metric called the “time to the first” one of these things. The process of converting an input into these things is commonly done using byte-pair encoding. During the prefill stage, (10[1])IDs corresponding to these things are embedded into a vector that is then passed through multiple transformer layers. The number of these things that can be processed at once is the context (10[1])window. (10[1])The cost of AI compute is often reported as (-5[1])the price per one million of these things. For 10 points, name these units of text that large language models split queries into and try to predict the next one of. ■END■ (0[1])

ANSWER: tokens [accept tokenization or tokenizers; accept time to first token; accept next token prediction or answers referring to trying to predict the next token; prompt on token-to-parameter ratio or tokens-per-parameter ratio; prompt on words or subwords or symbols or characters by asking “those are converted into what things?”]
<GC, Other Science> | Packet K - UCF A, Notre Dame B, Durham A, Michigan State A, Swarthmore A, Cambridge B
= Average correct buzzpoint

Back to tossups

Buzzes


Summary

TournamentEditionMatchHeardConv. %Neg %Avg. Buzz
California (South)Main Site4100%50%123.00
CanadaMain Site1283%25%111.30
Great LakesMain Site475%50%92.67
Lower Mid-AtlanticMain Site978%22%122.86
NortheastMain Site683%50%108.00
OverflowMain Site3100%33%125.33
SoutheastMain Site850%50%128.00
UKMain Site683%50%122.40
Upper Mid-AtlanticMain Site875%50%117.17