LLM Models & Open Source
45 regional LLMs · 185 open-source projects & datasets

Regional LLMs

+ Submit a model or project
SEA-LION v3 / v3.5
MIT / Apache-2.0 3B, 8B, 9B (text); 4B, 8B, 27B (vision, v4)
🇸🇬 Singapore · AI Singapore (AISG) · v3.5 (Llama-SEA-LION-v3-8B-IT, Gemma-SEA-LION-v3-9B-IT); v4 VLMs in preview
Southeast Asian Languages in One Network — the flagship open-source SEA LLM family from Singapore's national AI programme. Continued pre-training on ~500B tokens across 11 SEA languages with full instruction-tuning, alignment, and model-merging pipeline. Companion tools: SEA-HELM (eval benchmark) and SEA-Guard (safety).
Languages: enzhidvimsthmylotltakm
→ GitHub / Paper ↗
Sahabat-AI (Merak)
Apache-2.0 8B, 70B
🇮🇩 Indonesia · GoTo Group / GDP Labs · v1 (2024–2025)
Indonesia's sovereign LLM initiative by GoTo Group (Gojek/Tokopedia). Built on Llama 3 with continued pre-training and instruction fine-tuning on Indonesian-language data. Aimed at enterprise and government deployments in Indonesia, with a focus on Bahasa Indonesia fluency and local cultural context.
Languages: iden
→ GitHub / Paper ↗
SEA-LION v4 (Gemma)
Apache 2.0 27B
🇸🇬 Singapore · AI Singapore (AISG) · v4 (Gemma 3 27B, Nov 2025)
Flagship open multimodal SEA LLM, continually pre-trained on Gemma 3 27B with ~500B tokens across 11 SEA languages. Supports vision–language tasks, long-context understanding, and native function calling. Developed with Google Cloud VTC infrastructure.
Languages: enidkmlomsmytathtlvizh
→ GitHub / Paper ↗
SEA-LION v4 (Qwen)
Apache 2.0 32B
🇸🇬 Singapore · AI Singapore (AISG) × Alibaba Cloud · v4 (Qwen3 32B, Nov 2025)
Tops the SEA-HELM open-source leaderboard (under 200B params). Built on Qwen3-32B, continually pre-trained on ~100B tokens across 7 SEA languages. Efficient enough to run on a 32 GB consumer laptop.
Languages: enidmsmytathtlvi
→ GitHub / Paper ↗
Typhoon 2.5
Apache 2.0 Multiple sizes (3B–70B)
🇹🇭 Thailand · SCB 10X / OpenTyphoon · 2.5 (Oct 2025)
Latest open-source Thai LLM milestone with agentic design (tool use, multi-step reasoning), improved Thai fluency and tone, and high throughput efficiency. Previous Typhoon 2 offered 5 model sizes with multimodal (image + audio) variants.
Languages: then
→ GitHub / Paper ↗
Typhoon 2 / 2.1
Apache 2.0 7B / 8B / 70B
🇹🇭 Thailand · SCB 10X / OpenTyphoon · 2.1 (early 2025)
Production-grade Thai LLM series. Typhoon 1.5X featured a 70B variant rivalling leading models. Typhoon 2 added multimodal (image and audio) processing. Offers a free API via opentyphoon.ai and AWS-native deployment.
Languages: then
→ GitHub / Paper ↗
SeaLLM 3
Apache 2.0 7B / 70B
🇸🇬 Singapore · Alibaba DAMO Academy · SeaLLM 3 (2025)
Multilingual SEA LLM series from Alibaba DAMO Academy, built via continual pre-training on SEA language corpora, targeting linguistic and cultural alignment across the region. One of the earliest dedicated regional multilingual LLM initiatives.
Languages: enzhidthvimskmlomytl
→ GitHub / Paper ↗
WangchanX / WangchanLion
Apache 2.0 7B / 13B
🇹🇭 Thailand · NECTEC / VISTEC (Thailand) · WangchanLion (2024–2025)
Open Thai LLM series from Thailand's national electronics and computer technology centre (NECTEC) and VISTEC. Fine-tuned on Thai instruction datasets. WangchanX is the broader initiative covering Thai NLP tooling and benchmarks.
Languages: then
→ GitHub / Paper ↗
SEA-LION v4 (Qwen-SEA-LION-v4)
Apache 2.0 32B
🇸🇬 Singapore · AI Singapore (AISG) · v4 (Nov 2025)
Flagship open SEA LLM built on Qwen 32B; tops SEA-HELM leaderboard for open instruct models <200B. Pre-trained on 100B regional tokens from SEA-PILE-v2. 32k context window. Multimodal & reasoning-capable. Companion models: SEA-Guard (safety, Feb 2026) & SEA-LION-Embedding (Mar 2026).
Languages: enidmsthvitlmykmlotajv
→ GitHub / Paper ↗
Typhoon2.5
Apache 2.0 4B / 30B A3B (MoE)
🇹🇭 Thailand · SCB 10X (SCBX Group) · Typhoon 2.5 (2025–2026)
Latest generation of Thailand's premier Thai-English bilingual LLM family, built on Qwen3. Includes Typhoon2.5 4B and Typhoon2.5 30B A3B (MoE, 3B active). Family also includes reasoning model T1, multimodal, and translation-specific variants. Highest-performing Thai LLM on ThaiExam & M3Exam benchmarks.
Languages: then
→ GitHub / Paper ↗
Typhoon 2.1
Apache 2.0 4B / 12B
🇹🇭 Thailand · SCB 10X (SCBX Group) · Typhoon 2.1 (2025)
Thai-English bilingual LLM built on Gemma 3 base. Available in 4B and 12B sizes. Part of the broader Typhoon 2 family which also includes 1B, 3B, 8B instruct variants and a dedicated Typhoon-Translate 4B model for Thai↔English translation.
Languages: then
→ GitHub / Paper ↗
ILMU LLM
Proprietary / Sovereign Undisclosed
🇲🇾 Malaysia · YTL AI (YTL Corporation) / MYDigital · v1 (August 2025)
Malaysia's government-backed sovereign LLM launched August 2025, tailored for Bahasa Malaysia and hosted within YTL's 500MW Green Data Centre in Malaysia. Developed in partnership with NVIDIA. Targets e-commerce, media, and telecommunications anchor use cases. Aligned with Malaysia's national AI roadmap.
Languages: msen
→ GitHub / Paper ↗
Khmer LLM (Angkor Intelligence)
Apache 2.0 7B–13B
🇰🇭 Cambodia · Angkor Intelligence · v1 (Q4 2025 target)
Cambodia's first dedicated open-source Khmer LLM, trained on 50M tokens of Khmer text. Developed by Angkor Intelligence. A separate Khmer variant of SEA-LION 7B is being co-developed via an MoU between AI Forum Cambodia and AI Singapore (signed Jan 2025), marking the first official Khmer integration into the SEA-LION ecosystem.
Languages: kmen
→ GitHub / Paper ↗
SEA-LION v3
MIT / Apache 2.0 8B, 9B (also 3B; 70B reported)
🇸🇬 Singapore · AI Singapore (AISG) · v3 (Llama-SEA-LION-v3-8B-IT / Gemma-SEA-LION-v3-9B-IT)
Southeast Asian Languages in One Network — the flagship open-source SEA LLM family by AI Singapore, continued-pretrained on Llama 3 and Gemma 2 backbones over ~1T tokens spanning 11 SEA languages. Includes SEA-HELM evaluation benchmark and SEA-Guard safety layer.
Languages: enzhidvimsthmylofiltakm
→ GitHub / Paper ↗
Typhoon 2 / Typhoon 2.5
Apache 2.0 4B, 8B, 70B (multiple sizes incl. reasoning variant T1-3B)
🇹🇭 Thailand · SCB 10X · Typhoon 2.5 (2025–2026)
Thai-first bilingual LLM series by SCB 10X, optimized for Thai language understanding and instruction-following. Typhoon 2.5 is built on Qwen3; includes multimodal (vision + audio) variants and a Thai/English reasoning model (T1). Available via opentyphoon.ai API.
Languages: then
→ GitHub / Paper ↗
Typhoon / Typhoon 2.5
Apache 2.0 (base weights); API terms for hosted service 7B / 8B / 70B (multiple sizes across Typhoon 2.x family)
🇹🇭 Thailand · SCB 10X (SCBX Group) · Typhoon 2.5 (2025–2026)
Thailand's leading open-source LLM series optimised for Thai language, culture, and tokenisation. Typhoon 2.x adds multimodal (Typhoon Vision), audio (Typhoon Audio), Thai–English translation, and speech-to-text (incl. Isan dialect). Used in Thai government services and enterprise deployments.
Languages: then
→ GitHub / Paper ↗
GreenMind-Medium-14B-R1
Open-source (Apache 2.0 / check HuggingFace card) 14B
🇻🇳 Vietnam · GreenNode (VNG subsidiary) · 14B-R1 (Sept 2025)
The first open-source Vietnamese reasoning LLM, released by GreenNode (VNG's AI-cloud unit) in September 2025. Packaged on NVIDIA NIM and deployable on a single H100 GPU. Positioned as Vietnam's post-VinAI sovereign reasoning model.
Languages: vien
→ GitHub / Paper ↗
SeaLLMs
Apache 2.0 (Community License for >100M MAU) 7B
🇸🇬 Singapore (research base) · Alibaba DAMO Academy / Alibaba Group · SeaLLMs v3 (2024)
Southeast Asian Large Language Models by Alibaba DAMO Academy, released as part of the SEA open-model ecosystem. SeaLLMs v3 (7B) supports a broad suite of SEA languages via continued pre-training and instruction tuning. One of the earliest (Dec 2023) SEA-focused open LLMs alongside SEA-LION, covering Burmese, Khmer, Lao, Malay, Thai, Vietnamese, Indonesian, and more.
Languages: enzhidthvimsmykmlotl
→ GitHub / Paper ↗
GreenMind
Apache 2.0 14B
🇻🇳 Vietnam · GreenNode (VNG subsidiary) · Medium-14B-R1
First open-source Vietnamese reasoning LLM, released September 2025. Packaged on NVIDIA NIM for single-H100 deployability. Represents Vietnam's next-generation sovereign AI model after PhoGPT's freeze, with built-in chain-of-thought reasoning for Vietnamese.
Languages: vien
→ GitHub / Paper ↗
SEA-LION
Gemma Terms of Use (v4/v4.5); Apache 2.0 (v3) 27B (v4 flagship); 8B / 9B (v3 variants)
🇸🇬 Singapore · AI Singapore (AISG) · v4.5 (May 2026)
Southeast Asia's first family of open-source, multilingual, multimodal LLMs. v4 is built on Gemma 3 27B with 128K context, image+text understanding, function calling, and SEA-focused post-training. v4.5 expands agentic capabilities with a speculative decoder for up to 6× efficiency.
Languages: EnglishIndonesianMalayThaiVietnameseBurmeseTagalogKhmerLaoTamilMandarin
→ GitHub / Paper ↗
Typhoon
Apache 2.0 7B – 70B (multiple sizes across v1–v2.5)
🇹🇭 Thailand · SCB 10X · Typhoon 2.5
Thailand's flagship Thai-optimised LLM series. Typhoon 2.5 focuses on agentic AI with multi-step reasoning, improved function calling, high-throughput inference (3,000+ tokens/sec on H100), and superior Thai fluency. Multimodal variants (audio + vision) also available.
Languages: ThaiEnglish
→ GitHub / Paper ↗
SEA-LION v4.5
MIT / Gemma License (varies by variant) 27B (flagship Gemma-3-based)
🇸🇬 Singapore · AI Singapore (AISG) · v4.5
Southeast Asian Languages In One Network — a family of open-source LLMs fine-tuned for SEA languages and cultures. v4 (late 2025) introduced multimodality on Gemma 3 27B; v4.5 (March 2026) adds agent capabilities, a speculative decoder for 6× throughput, SEA-Guard safety models, and SEA-LION-Embedding suite. Supports 11 SEA languages. Part of Singapore's National Multi-Modal LLM Project.
Languages: enzhidvimsthmylofiltakm
→ GitHub / Paper ↗
Typhoon 2 / 2.5
Apache 2.0 (open weights) Multiple: 7B, 8B, 70B (text); multimodal & audio variants
🇹🇭 Thailand · SCB 10X / SCBX Group · Typhoon 2.5
Thailand's leading open-source Thai LLM series. Typhoon 2 features 5 model sizes, multimodal (image + audio) capabilities, and state-of-the-art Thai instruction-following. Typhoon 2.1 Gemma is featured in Google DeepMind's Gemmaverse. Typhoon 2.5 is available via opentyphoon.ai API. Deployed in Thai government services (OPDC partnership).
Languages: then
→ GitHub / Paper ↗
Sahabat-AI
Open (freely downloadable via Hugging Face) 70B (upgraded from 8B/9B launch)
🇮🇩 Indonesia · GoTo Group & Indosat Ooredoo Hutchison · 70B (June 2025)
Indonesia's sovereign open-source LLM, built on SEA-LION with AISG support. Upgraded to 70B parameters in June 2025 with multilingual chat service. Operates across Bahasa Indonesia and 4 local languages. All data and GPU infrastructure stored within Indonesian territory (GPU Merdeka sovereign cloud). Available via sahabat-ai.com and GoPay app.
Languages: idjvsubanbbcen
→ GitHub / Paper ↗
PhoGPT-4B / PhoGPT-4B-Chat
Open (research use) 4B (~3.7B actual)
🇻🇳 Vietnam · VinAI Research · PhoGPT-4B (2023, latest public release)
Vietnam's pioneering open-source generative LLM, pre-trained from scratch on a 102B-token Vietnamese corpus (482 GB cleaned). PhoGPT-4B-Chat is fine-tuned on 70K instruction prompts and 290K conversations. Uses a custom byte-level BPE tokenizer with 20K vocabulary types tailored for Vietnamese morphology.
Languages: vi
→ GitHub / Paper ↗
SEA-LION v4
MIT Multimodal
🇸🇬 Singapore · AI Singapore (AISG) · v4 (2026)
SEA-LION v4 adds multimodal capability (vision + language) to the family. Reasoning enhanced in v3.5. Part of Singapore's S$70M National Multimodal LLM Programme. Full SEA-LION ecosystem: base LLM, embeddings, safety guard.
Languages: IndonesianThaiVietnameseFilipinoBurmeseMalayLaoEnglishChineseKhmerTamil
→ GitHub / Paper ↗
SEA-LION v3.5
MIT 8B / 9B
🇸🇬 Singapore · AI Singapore (AISG) · v3.5 (Apr 2025)
SOTA on SEA-HELM multilingual benchmark. Continued pre-training on Llama-3.1-8B and Gemma-2-9B. 200B tokens, 16.8M instruction pairs. NVIDIA-optimized via TensorRT-LLM. MIT license. Widely deployed across ASEAN.
Languages: IndonesianThaiVietnameseFilipinoBurmeseMalayLaoEnglishChineseKhmerTamil
→ GitHub / Paper ↗
SEA-LION Embeddings
MIT Embedding model
🇸🇬 Singapore · AI Singapore (AISG) · 1.0 (Mar 2026)
SOTA retrieval embeddings for 10 SEA languages. Tested on SEA-BED (Southeast Asia Embedding Benchmark) with human-curated native data. Sets new records on retrieval, reranking, and semantic textual similarity. Essential for RAG in SEA languages.
Languages: IndonesianThaiVietnameseFilipinoMalayLaoEnglishChineseKhmerTamilBurmese
→ GitHub / Paper ↗
SEA-Guard
MIT Safety classifier
🇸🇬 Singapore · AI Singapore (AISG) · 1.0 (Feb 2026)
Dedicated safety layer for SEA-LION family. Culturally attuned — understands SEA-specific harmful content, cultural sensitivities, and local contexts. Launched Feb 4, 2026. Holistic NLP benchmarks + handcrafted SEA cultural diagnostic tests.
Languages: IndonesianThaiVietnameseFilipinoMalayEnglishChinese
→ GitHub / Paper ↗
MERaLiON-AudioLLM
Open weights 8B
🇸🇬 Singapore · A*STAR I2R + AI Singapore · 2 (2025)
Multimodal Empathetic Reasoning and Learning in One Network. Singapore's first AudioLLM. Trained on 62M multimodal samples, 260K hours of audio. SOTA on Singlish ASR. Fuses MERaLiON-Whisper encoder + SEA-LION text decoder. Free API trial available.
Languages: EnglishSinglishMalayChineseTamilThaiIndonesianVietnameseKhmerLaoBurmeseJavanese
→ GitHub / Paper ↗
Sailor2
Apache 2.0 1B / 8B / 20B
🇸🇬 Singapore · SEA AI Lab (Sea Limited) + SUTD · 2 (Dec 2024)
Built on Qwen2.5, continual pre-training on 500B tokens (400B SEA-specific). Sailor2-20B achieves 50-50 win rate vs GPT-4o across SEA languages. 13 SEA languages. Reproducibility cookbook published. Community-driven.
Languages: VietnameseThaiIndonesianMalayLaoEnglishChineseJavaneseSundaneseBurmeseTagalogKhmerTamil
→ GitHub / Paper ↗
SeaLLMs v3
Community 7B / 13B
🇸🇬 Singapore · Alibaba DAMO (Singapore lab) · 3 (2024)
Comprehensive SEA multilingual LLM. SeaLLM-13B outperforms ChatGPT-3.5 on SEA languages. Strong on low-resource languages like Lao and Khmer. SeaLLMs-Audio extension released Mar 2025.
Languages: IndonesianThaiVietnameseKhmerLaoMalayBurmeseTagalogEnglishChineseJavanese
→ GitHub / Paper ↗
SeaLLMs-Audio
Open weights 7B
🇸🇬 Singapore · Alibaba DAMO (Singapore lab) · 1 (Mar 2025)
First large audio-language model for SEA. Voice interactions across 5 languages. Built on Qwen2-Audio-7B. SOTA on SeaBench-Audio for Indonesian, Thai, and Vietnamese. Complements SeaLLMs text models.
Languages: IndonesianThaiVietnameseEnglishChinese
→ GitHub / Paper ↗
Typhoon2
Apache 2.0 7B / 8B
🇹🇭 Thailand · SCB 10X · 2 (2025)
Thai-English open LLM family from Siam Commercial Bank. Typhoon2-Instruct for chat, Typhoon2-Audio for Thai speech (SOTA). Collaboration with AI Singapore on cross-lingual audio. Most capable open Thai model.
Languages: ThaiEnglish
→ GitHub / Paper ↗
Pathumma
Government TBD
🇹🇭 Thailand · NSTDA Thailand · 1.0 (2025)
Thai sovereign LLM trained to understand Thai context and culture. Released by NSTDA. NVIDIA subsidiary investment for development. Government-owned, deployed across Thai public sector.
Languages: ThaiEnglish
→ GitHub / Paper ↗
OpenThaiGPT
Apache 2.0 7B / 13B
🇹🇭 Thailand · OpenThaiGPT Community · 1.0.0 (2024)
Open-source Thai GPT fine-tuned on Thai instruction datasets. Community-driven. Built on Llama and Mistral. Active development with Thai government and academic support. Most-used open Thai instruction model.
Languages: ThaiEnglish
→ GitHub / Paper ↗
WangchanBERTa
Apache 2.0 355M
🇹🇭 Thailand · NECTEC + AI Research Thailand · 1.0 (2021)
Pre-trained RoBERTa-based Thai language model. Trained on 78.5GB Thai text. Best-performing Thai BERT-class model. Used as backbone for Thai NLP tasks across industry and academia.
Languages: Thai
→ GitHub / Paper ↗
PhoGPT
Open 3.7B–4B
🇻🇳 Vietnam · VinAI Research · 4B (2023)
Vietnamese-first LLM pre-trained from scratch on 102B Vietnamese tokens. Avoids copyright issues by training on curated corpus. PhoGPT-4B-Chat fine-tuned on 300K+ conversations.
Languages: VietnameseEnglish
→ GitHub / Paper ↗
Vistral
Open 7B
🇻🇳 Vietnam · Viet AI Community · 7B-Chat (2024)
Vietnamese instruction-tuned LLM based on Mistral-7B. Community-built. Strong Vietnamese chat performance. Most downloaded open Vietnamese chat model on HuggingFace.
Languages: VietnameseEnglish
→ GitHub / Paper ↗
Vintern
Apache 2.0 1B / 3B
🇻🇳 Vietnam · VinAI Research · 1B / 3B (2025)
Compact Vietnamese multimodal LLMs (vision + language). VinAI's efficient models optimized for mobile and edge deployment. Strong performance on Vietnamese visual understanding tasks.
Languages: VietnameseEnglish
→ GitHub / Paper ↗
ILMU
Government TBD
🇲🇾 Malaysia · Malaysia NAIO · 1.0 (Aug 2025)
Malaysia's first domestically built sovereign LLM. Reflects Malaysian cultural values and constitutional principles. Launched Aug 2025 by National AI Office. Multilingual for Malaysia's four main language communities.
Languages: MalayEnglishTamilMandarin
→ GitHub / Paper ↗
MaLLaM
Open 1B / 3B / 5B
🇲🇾 Malaysia · Mesolitica · 1.0 (Jan 2024)
Family of Malay language models trained on 90B tokens from Malaysian contexts. First significant dedicated Malay LLM. Used in Malaysian government and enterprise. Mesolitica also maintains Malaya NLP library.
Languages: MalayEnglish
→ GitHub / Paper ↗
Merak
Apache 2.0 7B
🇮🇩 Indonesia · Koala Academy + community · 7B (2024)
Indonesian instruction-tuned LLM. Built on Llama-2-7B fine-tuned with Indonesian instruction datasets. Community-developed to address Indonesian language gap in major LLMs.
Languages: IndonesianEnglish
→ GitHub / Paper ↗
CendolBERT
MIT 355M
🇮🇩 Indonesia · IndoNLU Consortium · 2024
Multilingual BERT for Indonesian and regional languages. Covers Javanese and Sundanese — two of the world's most-spoken under-resourced languages. Used in IndoNLU benchmark evaluation.
Languages: IndonesianJavaneseSundanese
→ GitHub / Paper ↗
OpenSeal
Fully open 7B
🌏 ASEAN · Research Consortium · 2025
First fully open-source SEA LLM — all training data and code disclosed. Uses parallel data for continual pre-training of OLMo 2. Addresses LLM data transparency risks. Aimed at reproducibility.
Languages: IndonesianThaiVietnameseMalayLaoFilipinoEnglish
→ GitHub / Paper ↗