{"id":1224,"date":"2026-06-22T17:11:42","date_gmt":"2026-06-22T17:11:42","guid":{"rendered":"https:\/\/convly.ai\/model\/nvidia-nemotron-3-nano-omni\/"},"modified":"2026-08-03T02:30:25","modified_gmt":"2026-08-03T02:30:25","slug":"nvidia-nemotron-3-nano-omni","status":"publish","type":"ai_model","link":"https:\/\/convly.ai\/ar\/model\/nvidia-nemotron-3-nano-omni\/","title":{"rendered":"\u0625\u0646\u0641\u064a\u062f\u064a\u0627 \u0646\u064a\u0645\u0648\u062a\u0631\u0648\u0646 \u0663 \u0646\u0627\u0646\u0648 \u0623\u0648\u0645\u0646\u064a"},"content":{"rendered":"<h2>What is NVIDIA Nemotron 3 Nano Omni?<\/h2>\n<p>NVIDIA Nemotron 3 Nano Omni is an open omni-modal model: it sees, hears, watches and reads<br \/>\n\u2014 text, image, audio and video in, text out \u2014 from a single 30B-A3B mixture-of-experts that<br \/>\nactivates only about 3B parameters per token. It is a Mamba-Transformer hybrid with a 256K<br \/>\ncontext, released under the NVIDIA Open Model Agreement, which permits commercial use. It<br \/>\nscores 67.04 on OCRBench V2, 72.2 on Video-MME, 47.4 on OSWorld and 89.39 on Speech IF.<\/p>\n<p>Four input modalities in one 30B model that runs on a single RTX 5090 at NVFP4 (about<br \/>\n21 GB) is the notable part. The usual way to build an application that handles audio, video,<br \/>\nimages and text is to chain three or four specialist models, each with its own deployment,<br \/>\nlatency budget and failure mode. Nemotron 3 Nano Omni collapses that into one, which is why<br \/>\nits 47.4 on OSWorld \u2014 a computer-use benchmark that requires seeing a screen and acting on it<br \/>\n\u2014 is more interesting than the raw number suggests. There is no public per-token API price;<br \/>\nthe cost is the hardware, which for a single-GPU deployment is unusually approachable for a<br \/>\nmodel of this breadth.<\/p>","protected":false},"excerpt":{"rendered":"<p>What is NVIDIA Nemotron 3 Nano Omni? NVIDIA Nemotron 3 Nano Omni is an open omni-modal model: it sees, hears, [&hellip;]<\/p>\n","protected":false},"featured_media":0,"template":"","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"cm_developer":"NVIDIA","cm_model_type":"Multimodal (omni)","cm_modality":"Text, Image, Audio, Video \u2192 Text","cm_parameters":"30B total \/ ~3B active (MoE)","cm_context_window":"256K","cm_max_output":"\u2014","cm_license":"NVIDIA Open Model Agreement","cm_open_weights":"yes","cm_release_date":"2026","cm_input_price":"","cm_output_price":"","cm_api_providers":"Hugging Face, OpenRouter, NVIDIA NIM","cm_vram_fp16":"~62 GB","cm_vram_q4":"~21 GB (NVFP4)","cm_min_gpu":"RTX 5090 32GB (NVFP4) \/ H100 80GB (BF16)","cm_benchmarks":"OCRBench V2:67.04 | Video-MME:72.2 | OSWorld:47.4 | Speech IF:89.39","cm_official_url":"https:\/\/huggingface.co\/nvidia"},"class_list":["post-1224","ai_model","type-ai_model","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/convly.ai\/ar\/wp-json\/wp\/v2\/ai_model\/1224","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/convly.ai\/ar\/wp-json\/wp\/v2\/ai_model"}],"about":[{"href":"https:\/\/convly.ai\/ar\/wp-json\/wp\/v2\/types\/ai_model"}],"wp:attachment":[{"href":"https:\/\/convly.ai\/ar\/wp-json\/wp\/v2\/media?parent=1224"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}