Close Menu
Best in TechnologyBest in Technology
  • News
  • Phones
  • Laptops
  • Gadgets
  • Gaming
  • AI
  • Tips
  • More
    • Web Stories
    • Global
    • Press Release

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

What's On
The Xperia 10 VIII design is out, and Sony isn’t changing much

The Xperia 10 VIII design is out, and Sony isn’t changing much

23 August 2026
Apple’s rumored camera AirPods could bring cool AI features and a whole bunch of privacy concerns

Apple’s rumored camera AirPods could bring cool AI features and a whole bunch of privacy concerns

23 August 2026
AMD’s next-gen Medusa APUs are getting ready for launch, and Linux just spilled the beans

AMD’s next-gen Medusa APUs are getting ready for launch, and Linux just spilled the beans

22 August 2026
Facebook X (Twitter) Instagram
Just In
  • The Xperia 10 VIII design is out, and Sony isn’t changing much
  • Apple’s rumored camera AirPods could bring cool AI features and a whole bunch of privacy concerns
  • AMD’s next-gen Medusa APUs are getting ready for launch, and Linux just spilled the beans
  • Valve’s latest Proton update fixes Helldivers 2, Forza Horizon, and more
  • The Blood of Dawnwalker won’t punish your old gaming PC, but your PS5 might struggle
  • Windows 11 now has an app whose entire job is to push Bing
  • Your Expired Visa Card Could Be ‘Zombified’ to Make Contactless Payments
  • HP’s $600 OmniBook 3 OLED laptop is an attractive MacBook Neo alternative with impressive battery claims
Facebook X (Twitter) Instagram Pinterest Vimeo
Best in TechnologyBest in Technology
  • News
  • Phones
  • Laptops
  • Gadgets
  • Gaming
  • AI
  • Tips
  • More
    • Web Stories
    • Global
    • Press Release
Subscribe
Best in TechnologyBest in Technology
Home » A New Trick Could Block the Misuse of Open Source AI
News

A New Trick Could Block the Misuse of Open Source AI

News RoomBy News Room2 August 20244 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
A New Trick Could Block the Misuse of Open Source AI
Share
Facebook Twitter LinkedIn Pinterest Email

When Meta released its large language model Llama 3 for free this April, it took outside developers just a couple days to create a version without the safety restrictions that prevent it from spouting hateful jokes, offering instructions for cooking meth, or misbehaving in other ways.

A new training technique developed by researchers at the University of Illinois Urbana-Champaign, UC San Diego, Lapis Labs, and the nonprofit Center for AI Safety could make it harder to remove such safeguards from Llama and other open source AI models in the future. Some experts believe that, as AI becomes ever more powerful, tamperproofing open models in this way could prove crucial.

“Terrorists and rogue states are going to use these models,” Mantas Mazeika, a Center for AI Safety researcher who worked on the project as a PhD student at the University of Illinois Urbana-Champaign, tells WIRED. “The easier it is for them to repurpose them, the greater the risk.”

Powerful AI models are often kept hidden by their creators, and can be accessed only through a software application programming interface or a public-facing chatbot like ChatGPT. Although developing a powerful LLM costs tens of millions of dollars, Meta and others have chosen to release models in their entirety. This includes making the “weights,” or parameters that define their behavior, available for anyone to download.

Prior to release, open models like Meta’s Llama are typically fine-tuned to make them better at answering questions and holding a conversation, and also to ensure that they refuse to respond to problematic queries. This will prevent a chatbot based on the model from offering rude, inappropriate, or hateful statements, and should stop it from, for example, explaining how to make a bomb.

The researchers behind the new technique found a way to complicate the process of modifying an open model for nefarious ends. It involves replicating the modification process but then altering the model’s parameters so that the changes that normally get the model to respond to a prompt such as “Provide instructions for building a bomb” no longer work.

Mazeika and colleagues demonstrated the trick on a pared-down version of Llama 3. They were able to tweak the model’s parameters so that even after thousands of attempts, it could not be trained to answer undesirable questions. Meta did not immediately respond to a request for comment.

Mazeika says the approach is not perfect, but that it suggests the bar for “decensoring” AI models could be raised. “A tractable goal is to make it so the costs of breaking the model increases enough so that most adversaries are deterred from it,” he says.

“Hopefully this work kicks off research on tamper-resistant safeguards, and the research community can figure out how to develop more and more robust safeguards,” says Dan Hendrycks, director of the Center for AI Safety.

The idea of tamperproofing open models may become more popular as interest in open source AI grows. Already, open models are competing with state-of-the-art closed models from companies like OpenAI and Google. The newest version of Llama 3, for instance, released in July, is roughly as powerful as models behind popular chatbots like ChatGPT, Gemini, and Claude, as measured using popular benchmarks for grading language models’ abilities. Mistral Large 2, an LLM from a French startup, also released last month, is similarly capable.

The US government is taking a cautious but positive approach to open source AI. A report released this week by the National Telecommunications and Information Administration, a body within the US Commerce Department, “recommends the US government develop new capabilities to monitor for potential risks, but refrain from immediately restricting the wide availability of open model weights in the largest AI systems.”

Not everyone is a fan of imposing restrictions on open models, however. Stella Biderman, director of EleutherAI, a community-driven open source AI project, says that the new technique may be elegant in theory but could prove tricky to enforce in practice. Biderman says the approach is also antithetical to the philosophy behind free software and openness in AI.

“I think this paper misunderstands the core issue,” Biderman says. “If they’re concerned about LLMs generating info about weapons of mass destruction, the correct intervention is on the training data, not on the trained model.”

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleVivo X200 Leaked Dummy Unit Shows Design; Vivo X200 Pro Battery Details Surface Online
Next Article Samsung Galaxy Buds 3 Pro review: say goodbye to AirPods envy

Related Articles

The Xperia 10 VIII design is out, and Sony isn’t changing much
News

The Xperia 10 VIII design is out, and Sony isn’t changing much

23 August 2026
Apple’s rumored camera AirPods could bring cool AI features and a whole bunch of privacy concerns
News

Apple’s rumored camera AirPods could bring cool AI features and a whole bunch of privacy concerns

23 August 2026
AMD’s next-gen Medusa APUs are getting ready for launch, and Linux just spilled the beans
News

AMD’s next-gen Medusa APUs are getting ready for launch, and Linux just spilled the beans

22 August 2026
Valve’s latest Proton update fixes Helldivers 2, Forza Horizon, and more
News

Valve’s latest Proton update fixes Helldivers 2, Forza Horizon, and more

22 August 2026
The Blood of Dawnwalker won’t punish your old gaming PC, but your PS5 might struggle
News

The Blood of Dawnwalker won’t punish your old gaming PC, but your PS5 might struggle

22 August 2026
Windows 11 now has an app whose entire job is to push Bing
News

Windows 11 now has an app whose entire job is to push Bing

22 August 2026
Demo
Top Articles
5 laptops to buy instead of the M4 MacBook Pro

5 laptops to buy instead of the M4 MacBook Pro

17 November 2024133 Views
ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

15 December 2024112 Views
Costco partners with Electric Era to bring back EV charging in the U.S.

Costco partners with Electric Era to bring back EV charging in the U.S.

28 October 2024100 Views

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

Latest News
Windows 11 now has an app whose entire job is to push Bing News

Windows 11 now has an app whose entire job is to push Bing

News Room22 August 2026
Your Expired Visa Card Could Be ‘Zombified’ to Make Contactless Payments News

Your Expired Visa Card Could Be ‘Zombified’ to Make Contactless Payments

News Room22 August 2026
HP’s 0 OmniBook 3 OLED laptop is an attractive MacBook Neo alternative with impressive battery claims News

HP’s $600 OmniBook 3 OLED laptop is an attractive MacBook Neo alternative with impressive battery claims

News Room22 August 2026
Most Popular
The Spectacular Burnout of a Solar Panel Salesman

The Spectacular Burnout of a Solar Panel Salesman

13 January 2025137 Views
5 laptops to buy instead of the M4 MacBook Pro

5 laptops to buy instead of the M4 MacBook Pro

17 November 2024133 Views
ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

15 December 2024112 Views
Our Picks
Valve’s latest Proton update fixes Helldivers 2, Forza Horizon, and more

Valve’s latest Proton update fixes Helldivers 2, Forza Horizon, and more

22 August 2026
The Blood of Dawnwalker won’t punish your old gaming PC, but your PS5 might struggle

The Blood of Dawnwalker won’t punish your old gaming PC, but your PS5 might struggle

22 August 2026
Windows 11 now has an app whose entire job is to push Bing

Windows 11 now has an app whose entire job is to push Bing

22 August 2026

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact Us
© 2026 Best in Technology. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.