AISEAN / RESOURCES
LLM Models & Open Source
45 regional LLMs · 185 open-source projects & datasets
Regional LLMs
+ Submit a model or project
SEA-LION v3 / v3.5
MIT / Apache-2.0 3B, 8B, 9B (text); 4B, 8B, 27B (vision, v4)
Southeast Asian Languages in One Network — the flagship open-source SEA LLM family from Singapore's national AI programme. Continued pre-training on ~500B tokens across 11 SEA languages with full instruction-tuning, alignment, and model-merging pipeline. Companion tools: SEA-HELM (eval benchmark) and SEA-Guard (safety).
Languages: enzhidvimsthmylotltakm
Sahabat-AI (Merak)
Apache-2.0 8B, 70B
Indonesia's sovereign LLM initiative by GoTo Group (Gojek/Tokopedia). Built on Llama 3 with continued pre-training and instruction fine-tuning on Indonesian-language data. Aimed at enterprise and government deployments in Indonesia, with a focus on Bahasa Indonesia fluency and local cultural context.
Languages: iden
SEA-LION v4 (Gemma)
Apache 2.0 27B
Flagship open multimodal SEA LLM, continually pre-trained on Gemma 3 27B with ~500B tokens across 11 SEA languages. Supports vision–language tasks, long-context understanding, and native function calling. Developed with Google Cloud VTC infrastructure.
Languages: enidkmlomsmytathtlvizh
SEA-LION v4 (Qwen)
Apache 2.0 32B
Tops the SEA-HELM open-source leaderboard (under 200B params). Built on Qwen3-32B, continually pre-trained on ~100B tokens across 7 SEA languages. Efficient enough to run on a 32 GB consumer laptop.
Languages: enidmsmytathtlvi
Typhoon 2.5
Apache 2.0 Multiple sizes (3B–70B)
Latest open-source Thai LLM milestone with agentic design (tool use, multi-step reasoning), improved Thai fluency and tone, and high throughput efficiency. Previous Typhoon 2 offered 5 model sizes with multimodal (image + audio) variants.
Languages: then
Typhoon 2 / 2.1
Apache 2.0 7B / 8B / 70B
Production-grade Thai LLM series. Typhoon 1.5X featured a 70B variant rivalling leading models. Typhoon 2 added multimodal (image and audio) processing. Offers a free API via opentyphoon.ai and AWS-native deployment.
Languages: then
SeaLLM 3
Apache 2.0 7B / 70B
Multilingual SEA LLM series from Alibaba DAMO Academy, built via continual pre-training on SEA language corpora, targeting linguistic and cultural alignment across the region. One of the earliest dedicated regional multilingual LLM initiatives.
Languages: enzhidthvimskmlomytl
WangchanX / WangchanLion
Apache 2.0 7B / 13B
Open Thai LLM series from Thailand's national electronics and computer technology centre (NECTEC) and VISTEC. Fine-tuned on Thai instruction datasets. WangchanX is the broader initiative covering Thai NLP tooling and benchmarks.
Languages: then
SEA-LION v4 (Qwen-SEA-LION-v4)
Apache 2.0 32B
Flagship open SEA LLM built on Qwen 32B; tops SEA-HELM leaderboard for open instruct models <200B. Pre-trained on 100B regional tokens from SEA-PILE-v2. 32k context window. Multimodal & reasoning-capable. Companion models: SEA-Guard (safety, Feb 2026) & SEA-LION-Embedding (Mar 2026).
Languages: enidmsthvitlmykmlotajv
Typhoon2.5
Apache 2.0 4B / 30B A3B (MoE)
Latest generation of Thailand's premier Thai-English bilingual LLM family, built on Qwen3. Includes Typhoon2.5 4B and Typhoon2.5 30B A3B (MoE, 3B active). Family also includes reasoning model T1, multimodal, and translation-specific variants. Highest-performing Thai LLM on ThaiExam & M3Exam benchmarks.
Languages: then
Typhoon 2.1
Apache 2.0 4B / 12B
Thai-English bilingual LLM built on Gemma 3 base. Available in 4B and 12B sizes. Part of the broader Typhoon 2 family which also includes 1B, 3B, 8B instruct variants and a dedicated Typhoon-Translate 4B model for Thai↔English translation.
Languages: then
ILMU LLM
Proprietary / Sovereign Undisclosed
Malaysia's government-backed sovereign LLM launched August 2025, tailored for Bahasa Malaysia and hosted within YTL's 500MW Green Data Centre in Malaysia. Developed in partnership with NVIDIA. Targets e-commerce, media, and telecommunications anchor use cases. Aligned with Malaysia's national AI roadmap.
Languages: msen
Khmer LLM (Angkor Intelligence)
Apache 2.0 7B–13B
Cambodia's first dedicated open-source Khmer LLM, trained on 50M tokens of Khmer text. Developed by Angkor Intelligence. A separate Khmer variant of SEA-LION 7B is being co-developed via an MoU between AI Forum Cambodia and AI Singapore (signed Jan 2025), marking the first official Khmer integration into the SEA-LION ecosystem.
Languages: kmen
SEA-LION v3
MIT / Apache 2.0 8B, 9B (also 3B; 70B reported)
Southeast Asian Languages in One Network — the flagship open-source SEA LLM family by AI Singapore, continued-pretrained on Llama 3 and Gemma 2 backbones over ~1T tokens spanning 11 SEA languages. Includes SEA-HELM evaluation benchmark and SEA-Guard safety layer.
Languages: enzhidvimsthmylofiltakm
Typhoon 2 / Typhoon 2.5
Apache 2.0 4B, 8B, 70B (multiple sizes incl. reasoning variant T1-3B)
Thai-first bilingual LLM series by SCB 10X, optimized for Thai language understanding and instruction-following. Typhoon 2.5 is built on Qwen3; includes multimodal (vision + audio) variants and a Thai/English reasoning model (T1). Available via opentyphoon.ai API.
Languages: then
Typhoon / Typhoon 2.5
Apache 2.0 (base weights); API terms for hosted service 7B / 8B / 70B (multiple sizes across Typhoon 2.x family)
Thailand's leading open-source LLM series optimised for Thai language, culture, and tokenisation. Typhoon 2.x adds multimodal (Typhoon Vision), audio (Typhoon Audio), Thai–English translation, and speech-to-text (incl. Isan dialect). Used in Thai government services and enterprise deployments.
Languages: then
GreenMind-Medium-14B-R1
Open-source (Apache 2.0 / check HuggingFace card) 14B
The first open-source Vietnamese reasoning LLM, released by GreenNode (VNG's AI-cloud unit) in September 2025. Packaged on NVIDIA NIM and deployable on a single H100 GPU. Positioned as Vietnam's post-VinAI sovereign reasoning model.
Languages: vien
SeaLLMs
Apache 2.0 (Community License for >100M MAU) 7B
Southeast Asian Large Language Models by Alibaba DAMO Academy, released as part of the SEA open-model ecosystem. SeaLLMs v3 (7B) supports a broad suite of SEA languages via continued pre-training and instruction tuning. One of the earliest (Dec 2023) SEA-focused open LLMs alongside SEA-LION, covering Burmese, Khmer, Lao, Malay, Thai, Vietnamese, Indonesian, and more.
Languages: enzhidthvimsmykmlotl
GreenMind
Apache 2.0 14B
First open-source Vietnamese reasoning LLM, released September 2025. Packaged on NVIDIA NIM for single-H100 deployability. Represents Vietnam's next-generation sovereign AI model after PhoGPT's freeze, with built-in chain-of-thought reasoning for Vietnamese.
Languages: vien
SEA-LION
Gemma Terms of Use (v4/v4.5); Apache 2.0 (v3) 27B (v4 flagship); 8B / 9B (v3 variants)
Southeast Asia's first family of open-source, multilingual, multimodal LLMs. v4 is built on Gemma 3 27B with 128K context, image+text understanding, function calling, and SEA-focused post-training. v4.5 expands agentic capabilities with a speculative decoder for up to 6× efficiency.
Languages: EnglishIndonesianMalayThaiVietnameseBurmeseTagalogKhmerLaoTamilMandarin
Typhoon
Apache 2.0 7B – 70B (multiple sizes across v1–v2.5)
Thailand's flagship Thai-optimised LLM series. Typhoon 2.5 focuses on agentic AI with multi-step reasoning, improved function calling, high-throughput inference (3,000+ tokens/sec on H100), and superior Thai fluency. Multimodal variants (audio + vision) also available.
Languages: ThaiEnglish
SEA-LION v4.5
MIT / Gemma License (varies by variant) 27B (flagship Gemma-3-based)
Southeast Asian Languages In One Network — a family of open-source LLMs fine-tuned for SEA languages and cultures. v4 (late 2025) introduced multimodality on Gemma 3 27B; v4.5 (March 2026) adds agent capabilities, a speculative decoder for 6× throughput, SEA-Guard safety models, and SEA-LION-Embedding suite. Supports 11 SEA languages. Part of Singapore's National Multi-Modal LLM Project.
Languages: enzhidvimsthmylofiltakm
Typhoon 2 / 2.5
Apache 2.0 (open weights) Multiple: 7B, 8B, 70B (text); multimodal & audio variants
Thailand's leading open-source Thai LLM series. Typhoon 2 features 5 model sizes, multimodal (image + audio) capabilities, and state-of-the-art Thai instruction-following. Typhoon 2.1 Gemma is featured in Google DeepMind's Gemmaverse. Typhoon 2.5 is available via opentyphoon.ai API. Deployed in Thai government services (OPDC partnership).
Languages: then
Sahabat-AI
Open (freely downloadable via Hugging Face) 70B (upgraded from 8B/9B launch)
Indonesia's sovereign open-source LLM, built on SEA-LION with AISG support. Upgraded to 70B parameters in June 2025 with multilingual chat service. Operates across Bahasa Indonesia and 4 local languages. All data and GPU infrastructure stored within Indonesian territory (GPU Merdeka sovereign cloud). Available via sahabat-ai.com and GoPay app.
Languages: idjvsubanbbcen
PhoGPT-4B / PhoGPT-4B-Chat
Open (research use) 4B (~3.7B actual)
Vietnam's pioneering open-source generative LLM, pre-trained from scratch on a 102B-token Vietnamese corpus (482 GB cleaned). PhoGPT-4B-Chat is fine-tuned on 70K instruction prompts and 290K conversations. Uses a custom byte-level BPE tokenizer with 20K vocabulary types tailored for Vietnamese morphology.
Languages: vi
SEA-LION v4
MIT Multimodal
SEA-LION v4 adds multimodal capability (vision + language) to the family. Reasoning enhanced in v3.5. Part of Singapore's S$70M National Multimodal LLM Programme. Full SEA-LION ecosystem: base LLM, embeddings, safety guard.
Languages: IndonesianThaiVietnameseFilipinoBurmeseMalayLaoEnglishChineseKhmerTamil
SEA-LION v3.5
MIT 8B / 9B
SOTA on SEA-HELM multilingual benchmark. Continued pre-training on Llama-3.1-8B and Gemma-2-9B. 200B tokens, 16.8M instruction pairs. NVIDIA-optimized via TensorRT-LLM. MIT license. Widely deployed across ASEAN.
Languages: IndonesianThaiVietnameseFilipinoBurmeseMalayLaoEnglishChineseKhmerTamil
SEA-LION Embeddings
MIT Embedding model
SOTA retrieval embeddings for 10 SEA languages. Tested on SEA-BED (Southeast Asia Embedding Benchmark) with human-curated native data. Sets new records on retrieval, reranking, and semantic textual similarity. Essential for RAG in SEA languages.
Languages: IndonesianThaiVietnameseFilipinoMalayLaoEnglishChineseKhmerTamilBurmese
SEA-Guard
MIT Safety classifier
Dedicated safety layer for SEA-LION family. Culturally attuned — understands SEA-specific harmful content, cultural sensitivities, and local contexts. Launched Feb 4, 2026. Holistic NLP benchmarks + handcrafted SEA cultural diagnostic tests.
Languages: IndonesianThaiVietnameseFilipinoMalayEnglishChinese
MERaLiON-AudioLLM
Open weights 8B
Multimodal Empathetic Reasoning and Learning in One Network. Singapore's first AudioLLM. Trained on 62M multimodal samples, 260K hours of audio. SOTA on Singlish ASR. Fuses MERaLiON-Whisper encoder + SEA-LION text decoder. Free API trial available.
Languages: EnglishSinglishMalayChineseTamilThaiIndonesianVietnameseKhmerLaoBurmeseJavanese
Sailor2
Apache 2.0 1B / 8B / 20B
Built on Qwen2.5, continual pre-training on 500B tokens (400B SEA-specific). Sailor2-20B achieves 50-50 win rate vs GPT-4o across SEA languages. 13 SEA languages. Reproducibility cookbook published. Community-driven.
Languages: VietnameseThaiIndonesianMalayLaoEnglishChineseJavaneseSundaneseBurmeseTagalogKhmerTamil
SeaLLMs v3
Community 7B / 13B
Comprehensive SEA multilingual LLM. SeaLLM-13B outperforms ChatGPT-3.5 on SEA languages. Strong on low-resource languages like Lao and Khmer. SeaLLMs-Audio extension released Mar 2025.
Languages: IndonesianThaiVietnameseKhmerLaoMalayBurmeseTagalogEnglishChineseJavanese
SeaLLMs-Audio
Open weights 7B
First large audio-language model for SEA. Voice interactions across 5 languages. Built on Qwen2-Audio-7B. SOTA on SeaBench-Audio for Indonesian, Thai, and Vietnamese. Complements SeaLLMs text models.
Languages: IndonesianThaiVietnameseEnglishChinese
Typhoon2
Apache 2.0 7B / 8B
Thai-English open LLM family from Siam Commercial Bank. Typhoon2-Instruct for chat, Typhoon2-Audio for Thai speech (SOTA). Collaboration with AI Singapore on cross-lingual audio. Most capable open Thai model.
Languages: ThaiEnglish
Pathumma
Government TBD
Thai sovereign LLM trained to understand Thai context and culture. Released by NSTDA. NVIDIA subsidiary investment for development. Government-owned, deployed across Thai public sector.
Languages: ThaiEnglish
OpenThaiGPT
Apache 2.0 7B / 13B
Open-source Thai GPT fine-tuned on Thai instruction datasets. Community-driven. Built on Llama and Mistral. Active development with Thai government and academic support. Most-used open Thai instruction model.
Languages: ThaiEnglish
WangchanBERTa
Apache 2.0 355M
Pre-trained RoBERTa-based Thai language model. Trained on 78.5GB Thai text. Best-performing Thai BERT-class model. Used as backbone for Thai NLP tasks across industry and academia.
Languages: Thai
PhoGPT
Open 3.7B–4B
Vietnamese-first LLM pre-trained from scratch on 102B Vietnamese tokens. Avoids copyright issues by training on curated corpus. PhoGPT-4B-Chat fine-tuned on 300K+ conversations.
Languages: VietnameseEnglish
Vistral
Open 7B
Vietnamese instruction-tuned LLM based on Mistral-7B. Community-built. Strong Vietnamese chat performance. Most downloaded open Vietnamese chat model on HuggingFace.
Languages: VietnameseEnglish
Vintern
Apache 2.0 1B / 3B
Compact Vietnamese multimodal LLMs (vision + language). VinAI's efficient models optimized for mobile and edge deployment. Strong performance on Vietnamese visual understanding tasks.
Languages: VietnameseEnglish
ILMU
Government TBD
Malaysia's first domestically built sovereign LLM. Reflects Malaysian cultural values and constitutional principles. Launched Aug 2025 by National AI Office. Multilingual for Malaysia's four main language communities.
Languages: MalayEnglishTamilMandarin
MaLLaM
Open 1B / 3B / 5B
Family of Malay language models trained on 90B tokens from Malaysian contexts. First significant dedicated Malay LLM. Used in Malaysian government and enterprise. Mesolitica also maintains Malaya NLP library.
Languages: MalayEnglish
Merak
Apache 2.0 7B
Indonesian instruction-tuned LLM. Built on Llama-2-7B fine-tuned with Indonesian instruction datasets. Community-developed to address Indonesian language gap in major LLMs.
Languages: IndonesianEnglish
CendolBERT
MIT 355M
Multilingual BERT for Indonesian and regional languages. Covers Javanese and Sundanese — two of the world's most-spoken under-resourced languages. Used in IndoNLU benchmark evaluation.
Languages: IndonesianJavaneseSundanese
OpenSeal
Fully open 7B
First fully open-source SEA LLM — all training data and code disclosed. Uses parallel data for continual pre-training of OLMo 2. Addresses LLM data transparency risks. Aimed at reproducibility.
Languages: IndonesianThaiVietnameseMalayLaoFilipinoEnglish