A
aimodels44
Guest
gnani-evon-v3.3-30B-A3B is an English-and-Indic text-generation model from gnani, built for comprehension, extractive question answering, reasoning, instruction following, and tool-using applications. Its Mamba2–Transformer hybrid Mixture-of-Experts architecture uses the Nemotron Hybrid MoE (nemotron_h) network design, with 30B total parameters and about 3.5B active per token. It supports a 131,072-token (128K) context and BF16 precision. Gnani reports training it through continued pretraining, supervised fine-tuning, and GRPO reinforcement learning, with an English-and-Indic corpus and a focus on native-script Indic understanding. The key decision point is its specialization: benchmark results are strong for Indic comprehension and extractive QA, but mixed for translation, summarization, and instruction-following benchmarks. The model card lists Transformers, vLLM, and SGLang as runtimes; Transformers use requires trust_remote_code=True.Best use cases
Indic-language extractive question answering. For systems that retrieve passages and need answers grounded in those passages—for example, answering Hindi questions from policy documents or Bengali questions from a knowledge base—this model is a strong candidate to evaluate. Its reported XQuAD-in score is 64.59 F1 and XorQA-in score is 41.83 F1, leading the listed open-weight comparators on both. These are benchmark results, not a guarantee of performance on a particular corpus, so test retrieval quality, answer faithfulness, and script handling with your own data.
Indic comprehension and multilingual assistant chat. The model supports English, Hindi, Bengali, Telugu, Tamil, Marathi, Gujarati, Kannada, Malayalam, Odia, and Punjabi. Its MILU macro mean is 78.74 across those 11 languages, above the listed Sarvam-30B and Sarvam-105B scores. This makes it worth testing for multilingual assistants that need to understand questions across these languages; the Odia result is a notable exception, where Sarvam-105B scores higher.
Long-context multilingual RAG. The 131,072-token context window can accommodate long documents or larger retrieved context sets in a single request. The model card identifies multilingual RAG as an intended use. A long context limit does not establish retrieval accuracy or reliable use of every token, so measure answer quality and latency at the context lengths your application will send.
Multi-step reasoning in English and Indic languages. The model is intended for reasoning, and its training includes supervised capability data and GRPO reinforcement learning. The card reports a MILU Physics result of 91.32 versus 88.14 for
gpt-5.6-luna on a 4,435-item subject benchmark, and near parity with gpt-5.4-nano on a nine-language comprehension macro. For tasks that need multiple reasoning steps, the recommended configuration enables thinking and allows up to 16,384 generated tokens; validate accuracy and response cost on your own tasks.Tool-using assistants. The model is described as suitable for agent systems with tool calling and multi-turn reasoning. The vLLM example enables automatic tool choice and configures the
qwen3_coder tool-call parser and nano_v3 reasoning parser. This provides an integration path, but the README does not report tool-use benchmark scores or success rates, so test the exact tools, schemas, and failure recovery your agent requires.Limitations
The model card does not provide measured tokens per second, time-to-first-token, batch-size guidance, or a minimum VRAM figure. It lists NVIDIA H100 80GB, H200, and A100 hardware, but does not specify which configuration can serve the full 128K context or how many GPUs are needed. The 30B total parameter count and BF16 precision make hardware planning important; do not infer a specific memory requirement from the listed hardware alone.
The benchmark profile is uneven. The model leads the listed open-weight comparators on XQuAD-in and XorQA-in, but trails
gemma-4-26B-A4B and gemma-4-31B on MILU, IndicMMLU-Pro, and BharatMath. It also scores below the listed Sarvam models on CrossSum-in and IndicIFEval-ground, and far below the listed Gemma models on IndicIFEval-trans. Treat the model as a task-specific option, not a universal leader.The model card does not report a dedicated failure analysis, language-by-language safety results beyond Indicsafe5, or domain-specific reliability results. It explicitly advises against high-stakes decisions and legal, medical, or financial advice, and calls for safety evaluation in the target languages and domains. For strictly native-script output, the card recommends naming the language and script in a system instruction and checking the output script downstream.
The model uses BF16 in the provided examples. The README does not specify quantized checkpoints, quantization methods, model file format, or fine-tuning instructions. Transformers loading requires
trust_remote_code=True, which means deployment teams should review that code and their supply-chain controls before use. The license is Apache 2.0; consult the license terms for the intended deployment rather than treating the license label as a substitute for legal review.How it compares
sarvam-30b. Pick
gnani-evon-v3.3-30B-A3B when Indic comprehension or extractive QA is the priority: it scores 78.74 versus 67.15 on the MILU 11-language macro, 64.59 versus 37.86 on XQuAD-in, and 41.83 versus 22.86 on XorQA-in. Pick Sarvam-30B if your evaluation favors its results on tasks such as CrossSum-in or IndicIFEval-ground; the supplied comparison gives Sarvam-30B 9.66 versus 4.67 on CrossSum-in and 40.69 versus 28.66 on IndicIFEval-ground. Both are open-weight Indic-focused options in the supplied material, but the README provides no comparable speed or serving-cost measurements.deepseek-v3. The supplied description calls DeepSeek-V3-0324 a leading non-reasoning model and a milestone for open source, but provides no benchmark scores, hardware details, or cost figures that can be compared directly with this model. Choose
gnani-evon-v3.3-30B-A3B when the listed Indic-language coverage and its Indic QA results match your workload; choose DeepSeek-V3 when your own evaluation favors its general non-reasoning behavior. The available information does not support a concrete speed or quality ranking between them.nemotron-3-nano-omni. The supplied description identifies it as an open, efficient reasoning model for enterprise agentic workflows and gives it a 30B A3B hybrid Transformer–Mamba MoE design. That makes it architecturally related, but the provided information contains no benchmark results or runtime figures for a direct quality or speed comparison. Prefer
gnani-evon-v3.3-30B-A3B when its reported Indic-language benchmarks are relevant; evaluate Nemotron-3 Nano Omni when its stated enterprise-agent focus better matches the application.SraVaani-1.0. This is an ASR model, not a text-generation alternative: it is described as a roughly 430-million-parameter FastConformer model with a hybrid TDT-CTC decoder, trained for speech recognition across Indian languages and dialects. Choose it for speech-to-text; choose
gnani-evon-v3.3-30B-A3B for text chat, reasoning, and text-based RAG. Their tasks and output types differ, so speed and quality comparisons are not meaningful from the supplied data.sarvam-1-v0.5. The supplied description identifies this as an early checkpoint of
sarvam-2b, trained from scratch on 2 trillion tokens and intended for 10 Indic languages plus English; it also says a fully trained version is available under a different model name. The provided information does not include comparable benchmark results for this checkpoint. Consider gnani-evon-v3.3-30B-A3B when its reported Indic QA and comprehension scores matter; consider the smaller Sarvam checkpoint only after checking its current availability and evaluating it for your task. No speed or cost comparison is provided.Technical specifications
gnani-evon-v3.3-30B-A3B is a text-to-text, text-generation model. Its architecture combines Mamba2 and Transformer components with a Mixture-of-Experts design, using the Nemotron Hybrid MoE (nemotron_h) network architecture. It has 30B total parameters and approximately 3.5B active parameters per token. The model card specifies BF16 precision and a maximum context length of 131,072 tokens.Training comprised continued pretraining on a blended English-and-Indic corpus, supervised fine-tuning on instruction and capability-focused datasets covering QA, summarization, safety, and multi-turn dialogue in native scripts, and GRPO reinforcement learning aimed at reasoning quality, refusal behavior, and task completion on Indic-heavy prompts. Training ran on H200 GPU clusters using NVIDIA NeMo Curator, Megatron Bridge, and NeMo RL. Data collection and labeling used automated, human, and synthetic methods. The README does not give corpus size, token count, training steps, or total compute.
The model card lists Transformers 5.3.0, vLLM 0.12.0, and SGLang as runtime options, with Linux and NVIDIA H100 80GB, H200, or A100 hardware. The README specifies Transformers ≥ 5.3.0 and vLLM ≥ 0.12.0. It does not state a minimum GPU count, VRAM requirement, inference speed, batch-size limit, quantization option, or model file format.
Reported evaluation details include:
- MILU: 78.74 macro mean across 11 languages. Per-language scores: English 84.06, Kannada 82.44, Bengali 82.30, Hindi 81.99, Telugu 79.78, Gujarati 79.05, Tamil 78.12, Malayalam 77.00, Odia 66.94, Marathi 78.44, and Punjabi 76.02.
- Other Indic benchmarks: IndicMMLU-Pro accuracy 68.79; BharatMath accuracy 73.97; XQuAD-in F1 64.59; XorQA-in F1 41.83; FLORES-in chrF++ 34.52; CrossSum-in chrF 4.67; IndicIFEval-trans prompt-strict 47.45; IndicIFEval-ground prompt-strict 28.66; Indicsafe5 accuracy 86.50.
- Comparisons stated by the model card: MILU Physics 91.32 versus 88.14 for
gpt-5.6-lunaon 4,435 items; a nine-language Indic comprehension macro of 79.46 versus 79.69 forgpt-5.4-nano. The card reports a 0.23-point difference and says this model is ahead on four of those nine languages plus English. - License and usage: Apache 2.0; global deployment geography; not recommended for high-stakes decisions or legal, medical, or financial advice.
Model inputs and outputs
Inputs
- Text chat messages using OpenAI-compatible roles:
system,user,assistant, andtool. - Supported languages: English, Hindi, Bengali, Telugu, Tamil, Marathi, Gujarati, Kannada, Malayalam, Odia, and Punjabi.
- Maximum input/output size: 131,072 tokens. The model card describes this as a maximum context length; actual serving capacity depends on the runtime and hardware configuration.
- Chat templates can enable or disable reasoning traces. The README uses
enable_thinking=Falseto disable them.
Outputs
- Text strings, optionally including a reasoning trace.
- Tool-using applications can use the vLLM serving configuration with automatic tool choice and the specified tool-call and reasoning parsers.
- The model card recommends a system instruction naming the target language and script when native-script output is required, followed by downstream script validation.
Getting started
This Transformers example follows the model card. It loads the model in BF16, applies the chat template, and generates a Hindi response. Transformers loading requires
trust_remote_code=True.
Code:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "gnani/gnani-evon-v3.3"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
device_map="auto",
)
messages = [
{"role": "user", "content": "Explain photosynthesis in Hindi."},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=1024,
temperature=0.6,
top_p=0.95,
eos_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
The README recommends
temperature=0.6 and top_p=0.95 for reasoning, tool calling, and general instruction following. To disable reasoning traces, pass enable_thinking=False to apply_chat_template. For serving, the README also provides vLLM and SGLang examples. Its SGLang instructions recommend the Triton MoE backend to avoid a first-use CUTLASS compilation that requires matching CUDA compiler and header versions; they suggest --tp 2 or higher for longer context or higher concurrency and lowering --mem-fraction-static if graph capture runs out of memory.Frequently asked questions
Q: Can I use
gnani-evon-v3.3-30B-A3B commercially?A: The model card lists the Apache 2.0 license. Review the license terms and your organization’s requirements before deployment.
Q: What hardware or VRAM do I need to run it?
A: The README lists NVIDIA H100 80GB, H200, and A100 hardware on Linux. It does not specify minimum VRAM, GPU count, or the hardware needed to serve the full 128K context.
Q: How does it compare with Sarvam-30B for Indic question answering?
A: In the supplied results,
gnani-evon-v3.3-30B-A3B scores 64.59 versus 37.86 on XQuAD-in F1 and 41.83 versus 22.86 on XorQA-in F1. Sarvam-30B scores higher on CrossSum-in and IndicIFEval-ground in the same comparison.Q: What quality issues or failure modes are documented?
A: The README does not provide a detailed failure analysis. It reports weaker results than some listed models on translation, summarization, and instruction-following benchmarks, and recommends testing safety and output scripts in the target languages and domains.
Q: Can I fine-tune this model, and which framework supports it?
A: The README describes the model’s supervised fine-tuning and reinforcement-learning stages but does not provide fine-tuning
instructions or confirm a fine-tuning framework. It lists Transformers, vLLM, and SGLang as inference runtimes.
Q: What input format does the model expect?
A: It accepts text chat messages with OpenAI-compatible
system, user, assistant, and tool roles. The provided Transformers example formats messages with the tokenizer’s chat template.Q: How fast is inference, and what batch sizes are practical?
A: The README gives no tokens-per-second, latency, or batch-size measurements. Benchmark the intended runtime, context length, and hardware before estimating serving capacity.
Q: Is the model actively maintained?
A: The supplied information gives a release date of August 2026 but does not state a maintenance schedule or support policy.
Research context
The model card identifies the network architecture as Nemotron Hybrid MoE (
nemotron_h). Related research links provided for context include NVIDIA Nemotron Nano 2: Accurate and Efficient Hybrid and Nemotron 3 Ultra: Open, Efficient Mixture of Experts. These links provide architectural context; the supplied material does not claim that either paper evaluates this specific model.This is a simplified guide to an AI model called gnani-evon-v3.3-30B-A3B maintained by gnani.