Entrada de Pandipedia
Which tokenizer do gpt-oss models use?
The gpt-oss models utilize the o200k_harmony tokenizer, which is a Byte Pair Encoding (BPE) tokenizer. This tokenizer extends the o200k tokenizer used for other OpenAI models, such as GPT-4o and OpenAI o4-mini, and includes tokens specifically designed for the harmony chat format. The total number of tokens in this tokenizer is 201,088[1].
This tokenizer plays a crucial role in the models' training and processing capabilities, enabling effective communication in their agentic workflows and enhancing their instruction-following abilities[1].
Desa aquesta resposta
Crea el teu compte per conservar aquesta resposta i continuar-hi més tard.
Ho sentim, Pandi no ha pogut trobar una resposta.
Veiem alternatives:
- Modifica la consulta.
- Inicia un nou fil.
- Elimina les fonts (si s'han afegit manualment).