#comparativeAIandMLstudies
Exciting to develop @WorldUnivAndSch in all 200 countries https://t.co/DVguBYEiZ5 as world class #freeWUaSuniversities & in their 220 official languages (since India has 22 scheduled languages) for #comparativeAIandMLstudies too & in writing #SocInfoTechAndTheGlobalUniversity ~
— QuakerYogaMacFlower (@Q_YogaMacFlower) August 27, 2026
https://x.com/WUaSPress/
https://x.com/HarbinBook/
https://x.com/sgkmacleod/
https://x.com/TheOpenBand/
https://x.com/scottmacleod/sta
How could @WorldUnivAndSch w #ChinaWUaS #TibetWUaS & #NepalWUaS in dev #comparativeAIandMLstudies @ #freeWUaSuniversities & in writing #SocInfoTechAndTheGlobalUniversity
— WorldUnivandSch (@WorldUnivAndSch) August 28, 2026
study #AIandMLDisasterRelief eg in Xizang's mudslide-hit #GyirongChinahttps://t.co/d02mNLundp? #AIbenchmarks?
https://x.com/WorldUnivAndSch/
https://x.com/Q_YogaMacFlower/
https://x.com/WUaSPress/
https://x.com/HarbinBook/
https://x.com/sgkmacleod/
https://x.com/TheOpenBand/
https://x.com/scottmacleod/
#AIbenchmarks?
- Scott GK MacLeod
Founder, President, CEO & Professor
at / of best STEAM CC licensed OCW, Wiki,
World University & School (WUaS)
- USPS US Post Office, PO Box 132, General Delivery, Canyon, CA 94516
1) non-profit 501(c)(3) Public Charity
building on CC-4 licensed MIT OCW in 7 languages -
http://
2) for profit general stock company WUaS Corporation in CA - http://
- wuas_ceo@
https://wiki.
World University & School Innovation Research -
- http://scottmacleod.com
- https://twitter.com/
- https://x.com/Q_
- http://www.linkedin.com/in/
- https://twitter.com/
- https://scott-macleod.
- https://
* * *
Tracking foundational large language model (LLM) creation by country is an insightful lens, though counting them directly is tricky because "foundational" usually implies training a model from scratch—which requires vast compute, massive capital, and high-resource datasets—versus fine-tuning existing open-weights models (like Llama or Qwen).
Japan: ~10 to 20+ foundational models
Japan is an active developer of native, ground-up LLMs. Institutions like the National Institute of Informatics (LLM-jp), RIKEN, and private tech giants (NTT, CyberAgent, SoftBank, Fujitsu, and Preferred Networks) have trained dozens of proprietary and open-weights foundational models pre-trained heavily on Japanese and English text corpora.
Nigeria: 1 primary national model initiative
Nigeria recently spearheaded the creation of its first major sovereign multilingual LLM—developed in partnership with local AI firms (such as Awarri) and international research groups—pre-trained to cover several low-resource indigenous languages and accented English. Beyond this, most local projects focus on fine-tuning open global architectures rather than training models from scratch.
Cambodia: 0 to 1 emerging collaborative model
Cambodia does not currently have a fully native, independently trained foundational LLM. However, through partnerships like AI Forum Cambodia and AI Singapore, initiatives are underway to build and adapt specialized open-source Khmer LLMs to address regional language gaps.
Is this a useful metric?
Yes and no. Measuring by nation highlights digital sovereignty and linguistic representation, ensuring AI reflects local languages, laws, and cultures rather than relying entirely on foreign baselines. However, because foundational pre-training costs tens to hundreds of millions of dollars, most of the world relies on fine-tuning global open-weights models rather than building new ones from scratch
While Wikidata is not typically used to train full, standalone foundation models from scratch, it serves as a primary source for Knowledge-Graph Augmented LLMs (KG-LLMs), entity-aware translation, and Retrieval-Augmented Generation (RAG) frameworks globally.
Key AI Systems & Frameworks Using Wikidata:
KIF LLM Store (IBM): Integrates Wikidata to complete knowledge bases and ground LLM outputs with structured real-time facts.
Wikidata Embedding Project (Wikimedia Germany): Converts Wikidata's graph into vector embeddings so major AI models can interpret structured entities semantically.
Entity-Aware Translation Models: Research models (such as those featured in SemEval) use Wikidata IDs to help LLMs accurately translate culturally specific named entities across language pairs.
Cultural Evaluation Frameworks (e.g., MakiEval): Systems designed to benchmark multi-lingual LLM bias and global cultural awareness across 13+ languages using Wikidata as a cross-lingual anchor.
Countries Leading Wikidata-LLM Integration:
Germany: Home to Wikimedia Deutschland, which actively develops Wikidata vector integrations and embeddings for generative AI.
United States: Major research labs (IBM, Meta, academic institutions) leverage Wikidata for factual grounding, SPARQL-driven agent frameworks, and multi-lingual entity resolution.
Multinational European Collaborations: European NLP researchers heavily rely on Wikidata’s 348-language structure to reduce English-centric bias in localized AI models
*
https://en.wikipedia.org/wiki/Meconopsis
https://species.wikimedia.org/wiki/Papaveraceae
....
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.