Thursday, August 27, 2026

Himalayan Blue Poppy (Meconopsis species): #comparativeAIandMLstudies - Excited to develop @WorldUnivAndSch in all 200 countries https://wiki.worlduniversityandschool.org/wiki/Nation_States as world class #freeWUaSuniversities & in their 220 official languages (since India has 22 scheduled languages) for #comparativeAIandMLstudies too & in writing #SocInfoTechAndTheGlobalUniversity ~ * * * How could @WorldUnivAndSch w #ChinaWUaS #TibetWUaS & #NepalWUaS in dev #comparativeAIandMLstudies @ #freeWUaSuniversities & in writing #SocInfoTechAndTheGlobalUniversity study #AIandMLDisasterRelief eg in Xizang's mudslide-hit #GyirongChina https://news.cgtn.com/news/2026-08-28/Multiple-forces-mobilize-for-relief-in-mudslide-hit-Gyirong-1PYfsZ6erMA/p.html? #AIbenchmarks? * #AIbenchmarks? Platforms use of AI Foundational Large Language Models that countries create or ... * * * How many foundational large language models have, say the country of Cambodia, the country of Nigeria, the country of Japan created if that's a useful way of thinking about developing large language models that are foundational? * * * What large language models have been built and designed using Wikidata, Wikimedia's back-end structured knowledge database in 348 languages, if any, and in what countries especially?

 


#comparativeAIandMLstudies 


Excited to develop @WorldUnivAndSch in all 200 countries https://wiki.worlduniversityandschool.org/wiki/Nation_States as world class #freeWUaSuniversities & in their 220 official languages (since India has 22 scheduled languages) for #comparativeAIandMLstudies too & in writing #SocInfoTechAndTheGlobalUniversity ~





* * * 


How could @WorldUnivAndSch w #ChinaWUaS #TibetWUaS & #NepalWUaS in dev #comparativeAIandMLstudies @ #freeWUaSuniversities & in writing #SocInfoTechAndTheGlobalUniversity
study #AIandMLDisasterRelief eg in Xizang's mudslide-hit #GyirongChina





Multiple forces mobilized for disaster relief, supplies arrive in Xizang's mudslide-hit Gyirong
China
12:08, 28-Aug-2026





#AIbenchmarks?

Platforms use of AI

Foundational Large Language Models that countries create 

or ... .





--

Society, Information Technology, and the Global University, (forthcoming, Academic Press at World University and School, 2026) 

Scottish Small Piping album #2 - Honey Piobaireachd (2022)
Poetry! Order Book #5 Light, Float, Sit, Watsu ~ Virtually (2021, Academic Press at WUaS)

Scottish Small Piping album #1 Honey in the Bag ~ Out of the Air tune (2020)


Order Book #3 Winding Road Rainbow: Harbin, Wandering & the Poetry of Loving Bliss (2018,  Academic Press at WUaS)

Order Book #2 Haiku-ish and Other Loving Hippie Harbin Poetry (2017,  Academic Press at WUaS) 

Order Actual-Virtual Ethnographic Book #1: Naked Harbin Ethnography (2016, Academic Press at World University and School)



- Scott GK MacLeod  
Founder, President, CEO & Professor
at / of best STEAM CC licensed OCW, Wiki,
World University & School (WUaS)  
- USPS US Post Office, PO Box 132, General Delivery, Canyon, CA 94516

1) non-profit 501(c)(3) Public Charity
building on CC-4 licensed MIT OCW in 7 languages - 


2) for profit general stock company WUaS Corporation in CA - http://worlduniversityandschool.org/AcademicPress.html
wuas_ceo@worlduniversityandschool.org


https://wiki.worlduniversityandschool.org/wiki/Nation_States (planning ~200 countries' WUaS world class universities in their main languages, per the Olympics, for free-to-students' WUaS degrees from home)

https://wiki.worlduniversityandschool.org/wiki/Languages (planning to be in all 7159 living languages, each as or with wiki schools for open people-to-people wiki-teaching and wiki-learning, e.g. from here - https://wiki.worlduniversityandschool.org/wiki/Subjects)








* * * 

How many foundational large language models have, say the country of Cambodia, the country of Nigeria, the country of Japan created if that's a useful way of thinking about developing large language models that are foundational? 

Tracking foundational large language model (LLM) creation by country is an insightful lens, though counting them directly is tricky because "foundational" usually implies training a model from scratch—which requires vast compute, massive capital, and high-resource datasets—versus fine-tuning existing open-weights models (like Llama or Qwen).

  • Japan: ~10 to 20+ foundational models

    Japan is an active developer of native, ground-up LLMs. Institutions like the National Institute of Informatics (LLM-jp), RIKEN, and private tech giants (NTT, CyberAgent, SoftBank, Fujitsu, and Preferred Networks) have trained dozens of proprietary and open-weights foundational models pre-trained heavily on Japanese and English text corpora.

  • Nigeria: 1 primary national model initiative

    Nigeria recently spearheaded the creation of its first major sovereign multilingual LLM—developed in partnership with local AI firms (such as Awarri) and international research groups—pre-trained to cover several low-resource indigenous languages and accented English. Beyond this, most local projects focus on fine-tuning open global architectures rather than training models from scratch.

  • Cambodia: 0 to 1 emerging collaborative model

    Cambodia does not currently have a fully native, independently trained foundational LLM. However, through partnerships like AI Forum Cambodia and AI Singapore, initiatives are underway to build and adapt specialized open-source Khmer LLMs to address regional language gaps.

Is this a useful metric?

Yes and no. Measuring by nation highlights digital sovereignty and linguistic representation, ensuring AI reflects local languages, laws, and cultures rather than relying entirely on foreign baselines. However, because foundational pre-training costs tens to hundreds of millions of dollars, most of the world relies on fine-tuning global open-weights models rather than building new ones from scratch








* * *

What large language models have been built and designed using Wikidata, Wikimedia's back-end structured knowledge database in 348 languages, if any, and in what countries especially?

While Wikidata is not typically used to train full, standalone foundation models from scratch, it serves as a primary source for Knowledge-Graph Augmented LLMs (KG-LLMs), entity-aware translation, and Retrieval-Augmented Generation (RAG) frameworks globally.

  • Key AI Systems & Frameworks Using Wikidata:

    • KIF LLM Store (IBM): Integrates Wikidata to complete knowledge bases and ground LLM outputs with structured real-time facts.

    • Wikidata Embedding Project (Wikimedia Germany): Converts Wikidata's graph into vector embeddings so major AI models can interpret structured entities semantically.

    • Entity-Aware Translation Models: Research models (such as those featured in SemEval) use Wikidata IDs to help LLMs accurately translate culturally specific named entities across language pairs.

    • Cultural Evaluation Frameworks (e.g., MakiEval): Systems designed to benchmark multi-lingual LLM bias and global cultural awareness across 13+ languages using Wikidata as a cross-lingual anchor.

  • Countries Leading Wikidata-LLM Integration:

    • Germany: Home to Wikimedia Deutschland, which actively develops Wikidata vector integrations and embeddings for generative AI.

    • United States: Major research labs (IBM, Meta, academic institutions) leverage Wikidata for factual grounding, SPARQL-driven agent frameworks, and multi-lingual entity resolution.

    • Multinational European Collaborations: European NLP researchers heavily rely on Wikidata’s 348-language structure to reduce English-centric bias in localized AI models








https://upload.wikimedia.org/wikipedia/commons/thumb/6/6b/Himalayan_Blue_Poppy_%283638448884%29.jpg/1920px-Himalayan_Blue_Poppy_%283638448884%29.jpg?utm_source=commons.wikimedia.org&utm_campaign=index&utm_content=thumbnail&_=20130220023631














https://en.wikipedia.org/wiki/Meconopsis







https://species.wikimedia.org/wiki/Papaveraceae


....


No comments:

Post a Comment

Note: Only a member of this blog may post a comment.