Round 11: Tossup 18

The number of these things is the numerator of a ratio that is empirically optimal at about 20, the Chinchilla point. Computations involving the Q, K, and V vectors generated from these things are stored in a KV cache, which improves a metric called the “time to the first” one of these things. The process of converting an input into these things is commonly done using byte-pair encoding. During the prefill stage, IDs corresponding to these things are embedded into a vector that is then passed through multiple transformer layers. (-5[1])The number of these things that can be processed at once is the context window. The cost of (10[1])AI compute is often reported as the price per one million of these things. For 10 points, (10[1])name these units of text that large language models split queries into and try to predict the next one of. ■END■ (10[1])

ANSWER: tokens [accept tokenization or tokenizers; accept time to first token; accept next token prediction or answers referring to trying to predict the next token; prompt on token-to-parameter ratio or tokens-per-parameter ratio; prompt on words or subwords or symbols or characters by asking “those are converted into what things?”]
<GC, Other Science> | Packet K - UCF A, Notre Dame B, Durham A, Michigan State A, Swarthmore A, Cambridge B
= Average correct buzzpoint

Back to tossups

Buzzes


Summary

TournamentEditionMatchHeardConv. %Neg %Avg. Buzz
California (South)Main Site4100%50%123.00
CanadaMain Site1283%25%111.30
Great LakesMain Site475%50%92.67
Lower Mid-AtlanticMain Site978%22%122.86
NortheastMain Site683%50%108.00
OverflowMain Site3100%33%125.33
SoutheastMain Site850%50%128.00
UKMain Site683%50%122.40
Upper Mid-AtlanticMain Site875%50%117.17