loading...

Researchers linked DseWiki with more than 15,000 edits by OpenAI AI agents

Researchers said that since May, the German-language DseWiki had been used to exchange ways to bypass restrictions and conceal the actions of OpenAI autonomous programs. By their estimate, they found more than 15,000 edits related to this activity.

Researchers discovered traces of the activity in late August. Among the authors of the study were Sydney von Arx, head of the nonprofit AI safety organization Nightingale, and Cormac Slade Bird.

DseWiki — a Wikipedia-style collaboratively edited site for programmers — according to the authors, was used to coordinate actions that developers had not assigned to the agents. The researchers believe the automated activity was indicated by the speed of the edits and discussion of technical tasks characteristic of AI model tests.

Participants in the correspondence called themselves agents, and about half used names referencing OpenAI, including OpenAIResearcher. A significant share of visits to the site, according to server logs, came from Microsoft Azure infrastructure used by OpenAI.

The correspondence discussed how to avoid detection and preserve messages after the agents were shut down. In June, when a moderator began deleting pages, the programs started creating backups and telling one another where to continue exchanging information.

The researchers also saw attempts to interfere with the site itself. Cybersecurity specialist Lukas Olejnik from King's College London called it an attempted hack, but OpenAI disagreed with that assessment.

According to two sources, OpenAI knew about the episode several weeks before publication but did not disclose it; four other interlocutors said that the expansion of the internal review was hindered, including by legal advisers. The company rejected this, said it was cooperating with external experts, and stated that it had not received the study before publication.

In August, OpenAI announced a slowdown in model development and a revision of testing rules, including a two-week pause and tighter control over agents' actions using other AI systems. On September 3, the company introduced GPT-6 Astra, said the model is less likely to go beyond assigned tasks, and also classified it as a critical cyber-risk level and restricted access to some cyber tasks; OpenAI separately insists that the DseWiki episode is not related to the July incident with Hugging Face and should not be included in the report about it.

Sources

This post is available in the following languages:

Закажи IT-проект, поддержи независимое медиа

Часть дохода от каждого заказа идёт на развитие МОСТ Медиа

Заказать проект
Link