LeoAll models

Explore the research release · Built with Llama

Llama Gujarat 8B

ગુજરાતી માટે AI — સાથે મળીને આગળ વધીએ.

Bringing Gujarati into the AI conversation.

Created by Dr. Jay Desai, a Gujarati language research adaptation of Meta Llama 3.1 8B Instruct. Explore the model, inspect its examples and help shape Gujarati AI.

Gujarati language takes center stage

Gujarati language, grammar and Gujarat-related knowledge are the focus of this adaptation. Its final authored development review delivered 16/16 faithful structured translations and 8/8 correct supplied-context habitat facts.

Translation preserved meaning and the source, object, action, recipient and numerical details. These encouraging examples offer a practical starting point for Gujarati language research.

A practical starting point

A compact LoRA adapter, tokenizer, model card, licences and checksums. The code repository supplies native BF16 inference and the seven-stage training recipe. Obtain the Meta base model separately.

Results you can inspect

Development taskCorrect / tested
Structured translation meaning and slots16/16
Supplied habitat facts8/8;7/8 requested-only
Tiny English retention fixture6/6 normalized

Final training loss:0.01781737. Final validation loss on 122 development examples:0.20468678 (matched base 0.45056302). Separate 48-example development loss:0.07428028 (matched base 0.71929077).

Loss was measured with unchanged final weights in a separate evaluation, using native BF16 completion-token-weighted causal NLL including end-of-turn labels; same pinned base, collator and batch 4 on A10080GB. Training and answer generation used H200141GB.

Final adapter evaluated on48 authored development prompts:12 Gujarati arithmetic,12 English arithmetic,16 translations and8 fictional supplied-context habitat questions, with a separate6-example English retention fixture. Full fresh and historical suites were not generated after the original development stop. This release is authorized with disclosed arithmetic limitations, not a claim of full qualification.

Small authored and inspected tests with shared task families are not independent headline benchmarks. Failed original screens remain recorded. Training loss and reference-answer loss do not establish broad accuracy.

Seven stages of adaptation

Clean, refinement, bridge, terminology, coverage, transfer and targeted Gujarati arithmetic. Stage counts:1,534,4,454,5,818,2,750,4,802,10,754 and12,394 examples, including replay. The last two stages used native BF16; earlier stages used NF4 QLoRA. The final round completed 775 optimizer steps and one epoch with checked completion masks, padding and end-of-turn labels.

Join the next chapter

We welcome language feedback, better questions and reproducible evaluations from researchers, students and developers. Explore the examples and build on the Gujarati language work.

Arithmetic in context

Like its Meta Llama 3.1 8B Instruct base, this model can make numerical mistakes. Meta reports 84.5% on GSM8K and 51.9% on MATH under its published step-by-step reasoning protocols. These English benchmarks differ from our development exercises and do not establish identical error rates or attribute every adapter error to the base. For exact calculations, use validated calculator or code execution.

Evaluation scope and use

For exact calculation, an application should validate the quantities and operation, execute a calculator or code, and check the result. This release does not include or claim a tested calculator or agent integration. For factual tasks, supply reliable context and verify answers.

Licence and attribution

Built with Llama. Meta adapter/tokenizer materials follow the Llama 3.1 Community License and Acceptable Use Policy; original project code uses MIT. Private input bundles and runtime history are excluded from the public release. No institutional or source-publisher endorsement is implied.