• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Columns / Apertus: An LLM Designed for Data for Good

Apertus: An LLM Designed for Data for Good

October 1, 2025 Leave a Comment

an open book with binary code coming out of it
David Corliss

September 2 marked the release of a groundbreaking new large language model built with transparency and developing applications for public good as the guiding principle. Named for the Latin word for open, Apertus is the creation of the Swiss AI Initiative, an interdisciplinary team of AI experts from three leading Swiss research institutions: the Federal Technology Institute of Lausanne, Federal Institute of Technology in Zurich, and Swiss National Supercomputing Centre. In contrast to many commercial LLMs, Apertus is designed to be fully transparent and available for the public good, as well as promote digital sovereignty.

A key feature of the Apertus LLM is its multilingual design. Language limitations often constrain the utility of LLMs. For example, I am currently working with a group of researchers leveraging statistical science to fight human trafficking, with a lot of that work involving translating research into the language of the people serving victims and survivors. While trafficking is a problem all around the world, language limitations in conventional LLMs often limit their application in countries that could benefit the most. Apertus has been trained on 15 trillion tokens from more than 1,000 languages. Forty percent of the data driving this LMM comes from languages other than English, resulting in stronger support for underrepresented languages.

True to its name, Apertus raises the standard for openness and transparency. The entire pipeline, including the training data and source code, is publicly accessible. This supports reproducibility and encourages independent research and validation. It also makes it easier to demonstrate that practices comply with local laws and regulations in force where the data was collected, a principle known as data sovereignty.

The project comes from Switzerland, and it seems to me it carries something of a Swiss cultural world view in its design. Switzerland isn’t the kind of country we are familiar with in the United States, but rather a confederation of nearly independent cantons that see the value in working together on larger tasks serving a higher purpose. It is this perspective applied to LLM development that has led to Apertus, with this intentionally driving a greater ability to support Data for Good.

Getting Involved

In opportunities this month, the American Statistical Association’s Ethical Guidelines for Statistical Practice are due for a periodic review. The ASA’s Committee on Professional Ethics has issued a call for input to take into consideration when developing the update. This is your chance to be part of the process and have your voice heard. The committee is receiving input through the end of this year, will work on the update in 2026, and release the update in early 2027. Also, check out Stats + Stories, the ASA podcast featuring John Bailer and Rosemary Pennington from Miami University. This is one podcast you will want to follow.

I am a schoolteacher at heart and Apertus reminds me of something my favorite schoolteacher Shirley Chisholm used to say about herself: “unbought and unbossed.” Intentionally open and transparent, this LLM is designed with Data for Good in mind. It is a great example of how support for basic research in science and technology at a federal level, untethered to commercial demands and concerns, produces tools and applications that provide superior support for projects for the public good.

The project seeks to provide an AI infrastructure supporting public use, like roads, utilities, and schools. Apertus is issued using the Apache 2.0 open-source license, making it available to anyone. Researchers, government agencies, students, NGOs, and even commercial enterprises are all able to access and leverage this tool. Perhaps the most important thing about Apertus is the new pathway for LLM development it has created. It establishes a new process and standards for developing LLMs that serve the public interest.

Public outreach and access are facilitated through Hugging Face and GitHub. Small project proposals up to 50k GPU hours are accepted at any time, creating one avenue for getting started. One early project is democratizing LLMs for global languages; my colleagues on the translational team at the Global Association of Human Trafficking Scholars are thinking about writing a project proposal.

Just two years old, the Swiss AI Initiative, which created Apertus, has grown rapidly to become the largest open source/open science AI institute in the world. More than 800 staff members are supported by the Alps supercomputer with more than 10,000 GPUs. But for all its size, the initiative supports both small research projects and large ones. You can get all the details on their website, so check them out and see out what Apertus can bring to your next project in Data for Good.

A white man with a salt and pepper beard and hair smiles

David Corliss

With a PhD in statistical astrophysics, David Corliss works as a data scientist in industry. He serves on the ASA Board as a Council of Chapters representative and is the founder of Peace-Work, a data for good nongovernmental organization.

    This column is written for those interested in learning about the world of Data for Good, where statistical analysis is dedicated to good causes that benefit our lives, our communities, and our world. If you would like to know more or have ideas for articles, contact David Corliss.

    Filed Under: Columns, Stats4Good Tagged With: accessibility, AI, Apertus LLM, artificial intelligence, ASA, data, data for good, David Corliss, Education, GitHub, Global Association of Human Trafficking Scholars, Hugging Face, large language model, LLM, Research, statistician, statisticians, stats for good, stats4good, Swiss AI Initiative, transparency

    Reader Interactions

    Leave a Reply Cancel reply

    Your email address will not be published. Required fields are marked *

    Primary Sidebar

    Search

    More to See

    JSM 2027 Chicago, Illinois Logo

    Shape the Future: Invited Session Proposals Sought for JSM 2027

    August 3, 2026 By Meg Ruyle

    Laptop open next to the words "online community"

    How I Built a Portuguese-Language Analytics Community Online

    August 3, 2026 By Meg Ruyle

    New Member Spotlight: Collin Nill

    August 3, 2026 By Meg Ruyle

    colorful books

    Members Offer Advice for Writing as a Team

    August 3, 2026 By Meg Ruyle

    ASA HOME

    American Statistical Association

    Communications from the Executive Director

    ASA Leader Hub

    ASA Career Connect

    ADVERTISERS

    STATA
    SIAM

    Archives

    Categories

    Footer

    Editorial Staff

    Managing Editor
    Megan Murphy

    Graphic Designers / Production Coordinators
    Olivia Brown
    Meg Ruyle

    Communications Strategist
    Val Nirala

    Advertising Manager
    Christina Bonner

    Contributing Staff Members

    Kim Gilliam

    American Statistical Association
    277 South Washington Street, Suite 370
    Alexandria, VA 22314-3646
    Phone: (703) 302-1857

     

    Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in