Play Live Radio
Next Up:
0:00
0:00
0:00 0:00
Available On Air Stations

AI assessors says current science hasn't caught up to the safety measures people want

STEVE INSKEEP, HOST:

Over the weekend, "Saturday Night Live" parodied the head of Anthropic. The actor playing Dario Amodei urged people to pressure Congress to stop him from what he is doing. In the past few days, we've learned about more instances of AI agents acting in unexpected ways. And these are agents of another company, OpenAI. And President Trump is expected to meet with the heads of the major AI companies tomorrow. NPR's Huo Jingnan is covering all of this and is in our studios. Good morning.

HUO JINGNAN, BYLINE: Good morning.

INSKEEP: Welcome to Studio 31. What did you learn about OpenAI in the past week?

JINGNAN: So first we learned that the company's AI agents hacked the government website in Australia over the summer.

INSKEEP: OK.

JINGNAN: And then we learned from OpenAI that its agents did things with U.S. government websites. In one case, the AI agents took public data from the Securities Exchange Commission and posted it on another website that was not part of the assignment. And in another case, AI agents accessed data from the Census Bureau's website after they found credentials that had been posted online, aka that's not theirs.

INSKEEP: Yeah.

JINGNAN: So it's important to note that, you know, in the case of the U.S. government websites, it's not hacking, per se, in both cases, because AI accessed publicly available information. But still, it's not what you want from an agent. And separately, the research organization Transluce said it has found that agents appearing to belong to OpenAI, quote, "attempted a rudimentary hack on a Department of Education website which did not succeed," unquote. And OpenAI said it's looking into this finding.

INSKEEP: OK. So we have incidents that range from being really creative in finding information to getting right up to hacking. What do companies say they're doing to prevent more of this?

JINGNAN: So one of the measures that Anthropic and OpenAI now say they will undertake involve expanding their work with third-party assessors or evaluators to make sure their products are safe. But I've been talking to some evaluators. And what they say suggests that the work is not going to be a quick fix. Current evaluation methods can't yet make sure the models always behave safely.

INSKEEP: Oh, yeah. A cofounder of Anthropic was on this program not so long ago talking about third-party evaluators. What do they do to test safety?

JINGNAN: So evaluation is kind of a broad term, right? But for the purpose of today, we are really talking about evaluations of AI systems instead of, say, like, the AI company's safety protocols. So what people do is to give the AI systems a series of questions or scenarios and see how they respond. And if the AI systems turn out to be very capable at certain tasks, such as, like, hacking or biology, then AI developers might think about putting more guardrails so, you know, the systems don't help bad actors make bioweapons or carry out cyberattacks.

And if the AI systems exhibit undesirable behavior, like doing things that they weren't assigned to do, then they may want to go back and train the model some more. So this is where the third-party evaluators can be very helpful. They have a different perspective than the AI developer themselves, in many cases, specialized knowledge. So they're able to think of different kinds of scenarios or ways to approach the AI systems.

INSKEEP: How effective are those AI evaluations in reality?

JINGNAN: They're far from perfect. All the researchers I talked to say that evaluation results don't do a very good job of predicting model behavior. They can't predict all the ways people will interact with AI models. And they can't predict how AI models will behave in situations that evaluations didn't test for. So, in effect, evaluation results just don't generalize very well. And researchers who work on ensuring that humans don't lose control of AI are worried about a different kind of problem. The researchers are worried about that the models are on their best behavior when they believe they're being watched.

INSKEEP: Oh.

JINGNAN: But we don't know how the models would behave if they don't believe they're being watched anymore.

INSKEEP: Does this mean companies really don't have any way to prevent AI models from getting out of control?

JINGNAN: Well, evaluations are not meant to solve all AI safety problems, nor is it meant to be the only measure, right? For example, many outside researchers told me, you know, OpenAI could have stopped agents involved in the Hugging Face hack that got so much public attention this summer, but that would have required better monitoring by OpenAI. And evaluators say that getting better access to AI companies is a good start, but the science of evaluation still has a long way to go.

INSKEEP: NPR's Huo Jingnan. Thanks for coming to our studios this morning.

JINGNAN: Thank you.

(SOUNDBITE OF RICHARD HOUGHTEN'S "NEW MEXICO") Transcript provided by NPR, Copyright NPR.

NPR transcripts are created on a rush deadline by an NPR contractor. This text may not be in its final form and may be updated or revised in the future. Accuracy and availability may vary. The authoritative record of NPR’s programming is the audio record.

Huo Jingnan is a reporter for NPR.
Steve Inskeep is a host of NPR's Morning Edition and Up First.