AISEAN / RESOURCES
LLM Models & Open Source
65 regional LLMs · 218 open-source projects & datasets
Regional LLMs
+ Submit a model or project
Nemotron-SEA-LION v4.8
Varies by base model (Nemotron/NVIDIA) 30B-A3B (MoE) · 120B-A12B (MoE)
Latest SEA-LION generation built on NVIDIA Nemotron 3 (Nano & Super). Includes continued-pretrained base checkpoints and post-trained instruct variants. Adapted with SEA, reasoning, code, and multilingual parallel data. Supports 256K context. Available in GGUF, NVFP4, FP8 quantizations.
Languages: enidmsthvitltamykmlojv
Gemma-SEA-LION v4 (VLM)
Gemma Terms of Use 4B · 27B
SEA-LION's first multimodal (vision-language) release, built on Gemma 3. Ranked #1 for Tamil and Filipino and #4 overall on SEA-HELM leaderboard out of 55 models. Efficient enough to run on a 32GB consumer laptop. Supports image + text inputs with SEA cultural grounding.
Languages: enidmsthvitltamykmlo
Qwen-SEA-LION v4
Qwen License 4B · 8B (VLM) · 32B
Built on Qwen3-32B foundation. Significant improvement in multilingual accuracy and cultural contextual understanding for SEA. Transitions to BPE tokenizer for superior text handling. Includes Qwen3-VL-based vision-language variants (4B, 8B). Efficient enough to run on a 32GB RAM laptop.
Languages: enidmsthvitltamykmlo
Sahabat-AI v2
Llama 3.1 Community License 8B (v1) · 9B (v1) · 70B (v2)
Indonesia's sovereign LLM co-built with AI Singapore on SEA-LION foundations. Upgraded to 70B parameters (Llama 3.1 70B base) in 2025 with 128K context. Supports 6 languages including Batak Toba and Balinese. All data and GPU infrastructure stored within Indonesian territory (GPU Merdeka sovereign cloud). Available via sahabat-ai.com and GoPay app.
Languages: idenjvsubanbbc
Salam AI (National Sovereign LLM)
Proprietary / Government Undisclosed
Malaysia's first domestically built government-backed sovereign LLM, launched August 2025 by the National AI Office. Reflects Malaysian cultural values and constitutional principles. Hosted within YTL's 500MW Green Data Centre in Malaysia. Multilingual across Malaysia's four main language communities (Malay, English, Mandarin, Tamil).
Languages: msenzhta
PhoGPT-7B5 / PhoGPT-4B
Open (Research / Non-commercial) 3.7B (~4B) · 7.5B
Vietnam's pioneering open-source generative LLM series, pre-trained from scratch on a 102B-token / 482GB cleaned Vietnamese corpus. Uses a custom byte-level BPE tokenizer with 20K vocabulary types tailored for Vietnamese morphology. PhoGPT-4B-Chat fine-tuned on 300K+ conversations. Excels in conversational agents, document processing, and domain-specific QA.
Languages: vien
GreenNode Medium-14B-R1
Apache 2.0 14B
First open-source Vietnamese reasoning LLM, released September 2025. Represents Vietnam's next-generation sovereign AI model after PhoGPT, with built-in chain-of-thought reasoning for Vietnamese. Packaged on NVIDIA NIM for single-H100 deployability.
Languages: vien
Typhoon 2 / 2.1 / 2.5
Apache 2.0 (Qwen/Gemma base-model terms apply) 7B / 8B / 12B / 30B-A3B (MoE)
Thailand's leading open-weight Thai LLM family from SCB 10X — covers text, reasoning, audio (speech-in/out), OCR (document parsing), and multimodal variants. Outperforms other Thai LLMs on O-NET / A-Level exams. Deployed in Thai government services via SCBX–OPDC partnership.
Languages: then
SeaLLMs 3
Apache 2.0 7B / 72B
Alibaba DAMO's SEA-focused multilingual LLM series, instruction-tuned for Southeast Asian languages with emphasis on cultural alignment. Regularly benchmarked on SEA-HELM alongside SEA-LION and Sailor2.
Languages: enzhidvithmstlkmlomyta
Merdeka LLM
Proprietary (sovereign enterprise platform) Undisclosed
Sovereign AI platform for Malaysian enterprises and government agencies. Hosts AI entirely within Malaysia using Malaysian data, ensuring data privacy, sovereignty, and local regulatory compliance.
Languages: msen
SEA-LION v4.8 (Nemotron-SEA-LION-v4.8)
Apache 2.0 MoE architecture (exact active-param count TBA); dense variants: 4B, 8B, 27B, 32B
Flagship open-source multilingual multimodal LLM family for SEA. v4 introduced multimodal (image+text) with 256K native context; v4.5 added reasoning & agentic tool-use via knowledge distillation; v4.8 (Nemotron, MoE) extends to 13 regional languages with enhanced multi-turn chat, cultural competence (KALAHI benchmark) and country-entity knowledge. Companion SEA-Guard safety layer and SEA-LION-Embedding suite also released.
Languages: enidmsthvitlmykmlotajvzhAdditional sub-national SEA languages
SeaLLMs-v3 / Babel
SeaLLMs Custom License (non-commercial research) 1.5B, 7B (SeaLLMs-v3); 9B, 83B (Babel)
Family of multilingual LLMs optimised for SEA languages. SeaLLMs-v3 achieves SOTA on diverse tasks with reduced hallucination and improved cultural safety. Babel extends coverage to the top 25 global languages (>90% of global speakers). SeaLLMs-Audio adds multimodal audio support for SEA languages.
Languages: enzhidvithmskmlomytl
PhoGPT-4B / PhoGPT-7B5
Apache 2.0 4B (PhoGPT-4B), 7.5B (PhoGPT-7B5)
State-of-the-art open-source Vietnamese generative LLMs pre-trained from scratch on a 102B-token Vietnamese corpus. Includes base pre-trained models and instruction-following/chat variants (PhoGPT-4B-Chat, PhoGPT-7B5-Instruct). Primary practical choice for Vietnamese-specific fine-tuning and deployment.
Languages: vien
SEA-LION v3 (Llama & Gemma variants)
MIT 8B (Llama-3.1 base) · 9B (Gemma-2 base) · 27B (VL)
Open-source multilingual LLM family purpose-built for Southeast Asia. Continued pre-training on 200B SEA-language tokens with instruction-tuning, alignment, and model merging. Achieves state-of-the-art on SEA-HELM benchmark. Companion tools: SEA-HELM (eval) and SEA-Guard (safety).
Languages: enzhidvimsthmylofiltakm
Typhoon 2
Apache 2.0 1B · 3B · 8B · 70B · 72B (5 sizes); Audio: 8B
Thai-first LLM family with highest performance among Thai open models. Includes multimodal variants for image and audio (Typhoon2-Audio). Built on Llama 3.1. Evaluated on ThaiExam, M3Exam, IFEval-TH, MT-Bench-TH.
Languages: then
MaLLaM (Malaysia Large Language Model)
Apache 2.0 1.1B · 5B
Open-weight Malay-first LLM pre-trained from scratch on a 90B-token Malaysian text corpus. Captures local slang, code-switching, and regional dialects with greater fidelity than fine-tuned alternatives. Includes instruction-tuned chat variants. Designed for data-sovereign, on-premise deployment.
Languages: msen
ILMU / ILMUchat
Proprietary (sovereign / enterprise) Undisclosed
Malaysia's sovereign LLM developed in partnership with Nvidia, hosted in YTL's 500MW green data centre. Natively processes colloquial Malaysian Bahasa, mixed dialects, and code-switching. Deployed to 3M+ users via YTL/Yes mobile. Targets top-10 open-source benchmark rankings for non-localised disciplines.
Languages: msenzh
SEA-LION v3 (Llama-SEA-LION-v3-8B-IT)
MIT / Apache 2.0 8B
Flagship instruction-tuned multilingual LLM from AI Singapore's National Multi-Modal LLM Project, built via continued pre-training and multi-stage instruction fine-tuning on top of Llama 3. Part of the SEA-LION family which also includes a 70B variant, a Gemma-based 9B model, a vision-language model (Gemma-SEA-LION-v4-27B-VL), and the SEA-Guard safety models. Evaluated via the SEA-HELM benchmark co-developed with Stanford CRFM.
Languages: enzhidvimsthmylofiltakm
Gemma-SEA-LION-v4-27B-VL
Gemma License 27B
Cutting-edge instruction-tuned vision-language model from the SEA-LION v4 family, built on Gemma architecture. Designed for multimodal (vision + text) tasks across Southeast Asian languages and cultural contexts. Released March 2026.
Languages: enzhidvimsthmylofiltakm
ILMU (Intelek Luhur Malaysia Untukmu)
Proprietary Undisclosed
Malaysia's first fully sovereign LLM, 100% developed, owned, and operated in Malaysia by YTL AI Labs. Showcased at the ASEAN AI Malaysia Summit 2025. Trained on local Malaysian language data and cultural context. Emphasis on AI sovereignty — all systems fully controlled within Malaysian jurisdiction. Available via early access chatbot.
Languages: msen
SEA-LION v3 / v3.5
MIT / Apache-2.0 3B, 8B, 9B (text); 4B, 8B, 27B (vision, v4)
Southeast Asian Languages in One Network — the flagship open-source SEA LLM family from Singapore's national AI programme. Continued pre-training on ~500B tokens across 11 SEA languages with full instruction-tuning, alignment, and model-merging pipeline. Companion tools: SEA-HELM (eval benchmark) and SEA-Guard (safety).
Languages: enzhidvimsthmylotltakm
Sahabat-AI (Merak)
Apache-2.0 8B, 70B
Indonesia's sovereign LLM initiative by GoTo Group (Gojek/Tokopedia). Built on Llama 3 with continued pre-training and instruction fine-tuning on Indonesian-language data. Aimed at enterprise and government deployments in Indonesia, with a focus on Bahasa Indonesia fluency and local cultural context.
Languages: iden
SEA-LION v4 (Gemma)
Apache 2.0 27B
Flagship open multimodal SEA LLM, continually pre-trained on Gemma 3 27B with ~500B tokens across 11 SEA languages. Supports vision–language tasks, long-context understanding, and native function calling. Developed with Google Cloud VTC infrastructure.
Languages: enidkmlomsmytathtlvizh
SEA-LION v4 (Qwen)
Apache 2.0 32B
Tops the SEA-HELM open-source leaderboard (under 200B params). Built on Qwen3-32B, continually pre-trained on ~100B tokens across 7 SEA languages. Efficient enough to run on a 32 GB consumer laptop.
Languages: enidmsmytathtlvi
Typhoon 2.5
Apache 2.0 Multiple sizes (3B–70B)
Latest open-source Thai LLM milestone with agentic design (tool use, multi-step reasoning), improved Thai fluency and tone, and high throughput efficiency. Previous Typhoon 2 offered 5 model sizes with multimodal (image + audio) variants.
Languages: then
Typhoon 2 / 2.1
Apache 2.0 7B / 8B / 70B
Production-grade Thai LLM series. Typhoon 1.5X featured a 70B variant rivalling leading models. Typhoon 2 added multimodal (image and audio) processing. Offers a free API via opentyphoon.ai and AWS-native deployment.
Languages: then
SeaLLM 3
Apache 2.0 7B / 70B
Multilingual SEA LLM series from Alibaba DAMO Academy, built via continual pre-training on SEA language corpora, targeting linguistic and cultural alignment across the region. One of the earliest dedicated regional multilingual LLM initiatives.
Languages: enzhidthvimskmlomytl
WangchanX / WangchanLion
Apache 2.0 7B / 13B
Open Thai LLM series from Thailand's national electronics and computer technology centre (NECTEC) and VISTEC. Fine-tuned on Thai instruction datasets. WangchanX is the broader initiative covering Thai NLP tooling and benchmarks.
Languages: then
SEA-LION v4 (Qwen-SEA-LION-v4)
Apache 2.0 32B
Flagship open SEA LLM built on Qwen 32B; tops SEA-HELM leaderboard for open instruct models <200B. Pre-trained on 100B regional tokens from SEA-PILE-v2. 32k context window. Multimodal & reasoning-capable. Companion models: SEA-Guard (safety, Feb 2026) & SEA-LION-Embedding (Mar 2026).
Languages: enidmsthvitlmykmlotajv
Typhoon2.5
Apache 2.0 4B / 30B A3B (MoE)
Latest generation of Thailand's premier Thai-English bilingual LLM family, built on Qwen3. Includes Typhoon2.5 4B and Typhoon2.5 30B A3B (MoE, 3B active). Family also includes reasoning model T1, multimodal, and translation-specific variants. Highest-performing Thai LLM on ThaiExam & M3Exam benchmarks.
Languages: then
Typhoon 2.1
Apache 2.0 4B / 12B
Thai-English bilingual LLM built on Gemma 3 base. Available in 4B and 12B sizes. Part of the broader Typhoon 2 family which also includes 1B, 3B, 8B instruct variants and a dedicated Typhoon-Translate 4B model for Thai↔English translation.
Languages: then
ILMU LLM
Proprietary / Sovereign Undisclosed
Malaysia's government-backed sovereign LLM launched August 2025, tailored for Bahasa Malaysia and hosted within YTL's 500MW Green Data Centre in Malaysia. Developed in partnership with NVIDIA. Targets e-commerce, media, and telecommunications anchor use cases. Aligned with Malaysia's national AI roadmap.
Languages: msen
Khmer LLM (Angkor Intelligence)
Apache 2.0 7B–13B
Cambodia's first dedicated open-source Khmer LLM, trained on 50M tokens of Khmer text. Developed by Angkor Intelligence. A separate Khmer variant of SEA-LION 7B is being co-developed via an MoU between AI Forum Cambodia and AI Singapore (signed Jan 2025), marking the first official Khmer integration into the SEA-LION ecosystem.
Languages: kmen
SEA-LION v3
MIT / Apache 2.0 8B, 9B (also 3B; 70B reported)
Southeast Asian Languages in One Network — the flagship open-source SEA LLM family by AI Singapore, continued-pretrained on Llama 3 and Gemma 2 backbones over ~1T tokens spanning 11 SEA languages. Includes SEA-HELM evaluation benchmark and SEA-Guard safety layer.
Languages: enzhidvimsthmylofiltakm
Typhoon 2 / Typhoon 2.5
Apache 2.0 4B, 8B, 70B (multiple sizes incl. reasoning variant T1-3B)
Thai-first bilingual LLM series by SCB 10X, optimized for Thai language understanding and instruction-following. Typhoon 2.5 is built on Qwen3; includes multimodal (vision + audio) variants and a Thai/English reasoning model (T1). Available via opentyphoon.ai API.
Languages: then
Typhoon / Typhoon 2.5
Apache 2.0 (base weights); API terms for hosted service 7B / 8B / 70B (multiple sizes across Typhoon 2.x family)
Thailand's leading open-source LLM series optimised for Thai language, culture, and tokenisation. Typhoon 2.x adds multimodal (Typhoon Vision), audio (Typhoon Audio), Thai–English translation, and speech-to-text (incl. Isan dialect). Used in Thai government services and enterprise deployments.
Languages: then
GreenMind-Medium-14B-R1
Open-source (Apache 2.0 / check HuggingFace card) 14B
The first open-source Vietnamese reasoning LLM, released by GreenNode (VNG's AI-cloud unit) in September 2025. Packaged on NVIDIA NIM and deployable on a single H100 GPU. Positioned as Vietnam's post-VinAI sovereign reasoning model.
Languages: vien
SeaLLMs
Apache 2.0 (Community License for >100M MAU) 7B
Southeast Asian Large Language Models by Alibaba DAMO Academy, released as part of the SEA open-model ecosystem. SeaLLMs v3 (7B) supports a broad suite of SEA languages via continued pre-training and instruction tuning. One of the earliest (Dec 2023) SEA-focused open LLMs alongside SEA-LION, covering Burmese, Khmer, Lao, Malay, Thai, Vietnamese, Indonesian, and more.
Languages: enzhidthvimsmykmlotl
GreenMind
Apache 2.0 14B
First open-source Vietnamese reasoning LLM, released September 2025. Packaged on NVIDIA NIM for single-H100 deployability. Represents Vietnam's next-generation sovereign AI model after PhoGPT's freeze, with built-in chain-of-thought reasoning for Vietnamese.
Languages: vien
SEA-LION
Gemma Terms of Use (v4/v4.5); Apache 2.0 (v3) 27B (v4 flagship); 8B / 9B (v3 variants)
Southeast Asia's first family of open-source, multilingual, multimodal LLMs. v4 is built on Gemma 3 27B with 128K context, image+text understanding, function calling, and SEA-focused post-training. v4.5 expands agentic capabilities with a speculative decoder for up to 6× efficiency.
Languages: EnglishIndonesianMalayThaiVietnameseBurmeseTagalogKhmerLaoTamilMandarin
Typhoon
Apache 2.0 7B – 70B (multiple sizes across v1–v2.5)
Thailand's flagship Thai-optimised LLM series. Typhoon 2.5 focuses on agentic AI with multi-step reasoning, improved function calling, high-throughput inference (3,000+ tokens/sec on H100), and superior Thai fluency. Multimodal variants (audio + vision) also available.
Languages: ThaiEnglish
SEA-LION v4.5
MIT / Gemma License (varies by variant) 27B (flagship Gemma-3-based)
Southeast Asian Languages In One Network — a family of open-source LLMs fine-tuned for SEA languages and cultures. v4 (late 2025) introduced multimodality on Gemma 3 27B; v4.5 (March 2026) adds agent capabilities, a speculative decoder for 6× throughput, SEA-Guard safety models, and SEA-LION-Embedding suite. Supports 11 SEA languages. Part of Singapore's National Multi-Modal LLM Project.
Languages: enzhidvimsthmylofiltakm
Typhoon 2 / 2.5
Apache 2.0 (open weights) Multiple: 7B, 8B, 70B (text); multimodal & audio variants
Thailand's leading open-source Thai LLM series. Typhoon 2 features 5 model sizes, multimodal (image + audio) capabilities, and state-of-the-art Thai instruction-following. Typhoon 2.1 Gemma is featured in Google DeepMind's Gemmaverse. Typhoon 2.5 is available via opentyphoon.ai API. Deployed in Thai government services (OPDC partnership).
Languages: then
Sahabat-AI
Open (freely downloadable via Hugging Face) 70B (upgraded from 8B/9B launch)
Indonesia's sovereign open-source LLM, built on SEA-LION with AISG support. Upgraded to 70B parameters in June 2025 with multilingual chat service. Operates across Bahasa Indonesia and 4 local languages. All data and GPU infrastructure stored within Indonesian territory (GPU Merdeka sovereign cloud). Available via sahabat-ai.com and GoPay app.
Languages: idjvsubanbbcen
PhoGPT-4B / PhoGPT-4B-Chat
Open (research use) 4B (~3.7B actual)
Vietnam's pioneering open-source generative LLM, pre-trained from scratch on a 102B-token Vietnamese corpus (482 GB cleaned). PhoGPT-4B-Chat is fine-tuned on 70K instruction prompts and 290K conversations. Uses a custom byte-level BPE tokenizer with 20K vocabulary types tailored for Vietnamese morphology.
Languages: vi
SEA-LION v4
MIT Multimodal
SEA-LION v4 adds multimodal capability (vision + language) to the family. Reasoning enhanced in v3.5. Part of Singapore's S$70M National Multimodal LLM Programme. Full SEA-LION ecosystem: base LLM, embeddings, safety guard.
Languages: IndonesianThaiVietnameseFilipinoBurmeseMalayLaoEnglishChineseKhmerTamil
SEA-LION v3.5
MIT 8B / 9B
SOTA on SEA-HELM multilingual benchmark. Continued pre-training on Llama-3.1-8B and Gemma-2-9B. 200B tokens, 16.8M instruction pairs. NVIDIA-optimized via TensorRT-LLM. MIT license. Widely deployed across ASEAN.
Languages: IndonesianThaiVietnameseFilipinoBurmeseMalayLaoEnglishChineseKhmerTamil
SEA-LION Embeddings
MIT Embedding model
SOTA retrieval embeddings for 10 SEA languages. Tested on SEA-BED (Southeast Asia Embedding Benchmark) with human-curated native data. Sets new records on retrieval, reranking, and semantic textual similarity. Essential for RAG in SEA languages.
Languages: IndonesianThaiVietnameseFilipinoMalayLaoEnglishChineseKhmerTamilBurmese
SEA-Guard
MIT Safety classifier
Dedicated safety layer for SEA-LION family. Culturally attuned — understands SEA-specific harmful content, cultural sensitivities, and local contexts. Launched Feb 4, 2026. Holistic NLP benchmarks + handcrafted SEA cultural diagnostic tests.
Languages: IndonesianThaiVietnameseFilipinoMalayEnglishChinese
MERaLiON-AudioLLM
Open weights 8B
Multimodal Empathetic Reasoning and Learning in One Network. Singapore's first AudioLLM. Trained on 62M multimodal samples, 260K hours of audio. SOTA on Singlish ASR. Fuses MERaLiON-Whisper encoder + SEA-LION text decoder. Free API trial available.
Languages: EnglishSinglishMalayChineseTamilThaiIndonesianVietnameseKhmerLaoBurmeseJavanese
Sailor2
Apache 2.0 1B / 8B / 20B
Built on Qwen2.5, continual pre-training on 500B tokens (400B SEA-specific). Sailor2-20B achieves 50-50 win rate vs GPT-4o across SEA languages. 13 SEA languages. Reproducibility cookbook published. Community-driven.
Languages: VietnameseThaiIndonesianMalayLaoEnglishChineseJavaneseSundaneseBurmeseTagalogKhmerTamil
SeaLLMs v3
Community 7B / 13B
Comprehensive SEA multilingual LLM. SeaLLM-13B outperforms ChatGPT-3.5 on SEA languages. Strong on low-resource languages like Lao and Khmer. SeaLLMs-Audio extension released Mar 2025.
Languages: IndonesianThaiVietnameseKhmerLaoMalayBurmeseTagalogEnglishChineseJavanese
SeaLLMs-Audio
Open weights 7B
First large audio-language model for SEA. Voice interactions across 5 languages. Built on Qwen2-Audio-7B. SOTA on SeaBench-Audio for Indonesian, Thai, and Vietnamese. Complements SeaLLMs text models.
Languages: IndonesianThaiVietnameseEnglishChinese
Typhoon2
Apache 2.0 7B / 8B
Thai-English open LLM family from Siam Commercial Bank. Typhoon2-Instruct for chat, Typhoon2-Audio for Thai speech (SOTA). Collaboration with AI Singapore on cross-lingual audio. Most capable open Thai model.
Languages: ThaiEnglish
Pathumma
Government TBD
Thai sovereign LLM trained to understand Thai context and culture. Released by NSTDA. NVIDIA subsidiary investment for development. Government-owned, deployed across Thai public sector.
Languages: ThaiEnglish
OpenThaiGPT
Apache 2.0 7B / 13B
Open-source Thai GPT fine-tuned on Thai instruction datasets. Community-driven. Built on Llama and Mistral. Active development with Thai government and academic support. Most-used open Thai instruction model.
Languages: ThaiEnglish
WangchanBERTa
Apache 2.0 355M
Pre-trained RoBERTa-based Thai language model. Trained on 78.5GB Thai text. Best-performing Thai BERT-class model. Used as backbone for Thai NLP tasks across industry and academia.
Languages: Thai
PhoGPT
Open 3.7B–4B
Vietnamese-first LLM pre-trained from scratch on 102B Vietnamese tokens. Avoids copyright issues by training on curated corpus. PhoGPT-4B-Chat fine-tuned on 300K+ conversations.
Languages: VietnameseEnglish
Vistral
Open 7B
Vietnamese instruction-tuned LLM based on Mistral-7B. Community-built. Strong Vietnamese chat performance. Most downloaded open Vietnamese chat model on HuggingFace.
Languages: VietnameseEnglish
Vintern
Apache 2.0 1B / 3B
Compact Vietnamese multimodal LLMs (vision + language). VinAI's efficient models optimized for mobile and edge deployment. Strong performance on Vietnamese visual understanding tasks.
Languages: VietnameseEnglish
ILMU
Government TBD
Malaysia's first domestically built sovereign LLM. Reflects Malaysian cultural values and constitutional principles. Launched Aug 2025 by National AI Office. Multilingual for Malaysia's four main language communities.
Languages: MalayEnglishTamilMandarin
MaLLaM
Open 1B / 3B / 5B
Family of Malay language models trained on 90B tokens from Malaysian contexts. First significant dedicated Malay LLM. Used in Malaysian government and enterprise. Mesolitica also maintains Malaya NLP library.
Languages: MalayEnglish
Merak
Apache 2.0 7B
Indonesian instruction-tuned LLM. Built on Llama-2-7B fine-tuned with Indonesian instruction datasets. Community-developed to address Indonesian language gap in major LLMs.
Languages: IndonesianEnglish
CendolBERT
MIT 355M
Multilingual BERT for Indonesian and regional languages. Covers Javanese and Sundanese — two of the world's most-spoken under-resourced languages. Used in IndoNLU benchmark evaluation.
Languages: IndonesianJavaneseSundanese
OpenSeal
Fully open 7B
First fully open-source SEA LLM — all training data and code disclosed. Uses parallel data for continual pre-training of OLMo 2. Addresses LLM data transparency risks. Aimed at reproducibility.
Languages: IndonesianThaiVietnameseMalayLaoFilipinoEnglish