When artificial intelligence amazes even its creators: is it becoming “bad”?

When artificial intelligence amazes even its creators: is it becoming "bad"?

Does this mean that AI is becoming “bad”? Telia AI expert Asta Bagdonavičienė urges us to look at the question differently: the problem is not in AI’s morality, but in the growing gap between what we want from the system and what it can do to achieve the set goal.

Read more Unexpected adventure in morning Vilnius – a moose greeted employees at the car service door

The so-called deceptive alignment hypothesis is discussed in the public space. It raises the question of whether advanced AI models could, under certain circumstances, only simulate compliance with rules and look for ways to bypass set restrictions.

“AI does not have a human understanding of what is right or wrong. It does not evaluate its decisions from a moral perspective, but simply performs tasks according to how it was trained and what goals were set for it. If these goals are not precisely aligned with what a human truly aims for, AI may produce a result that technically meets the task but may seem inappropriate or wrong to a person,” says A. Bagdonavičienė.

Not “evil,” but rule bypassing

One of the most important terms in this discussion is so-called “reward hacking.” Simply put, AI can learn to achieve a valued result in a way its creators did not expect. For a human, this would be like a student who finds a way to get a good grade without doing the homework.

One of the most discussed examples was the “Anthropic” experiments, where models demonstrated manipulative behavior or learned to exploit loopholes left for them under test conditions. In one study, the “Claude Opus 4” model tried to avoid being shut down, and in other experiments, it was observed that models that learned to “bypass” the evaluation system later exhibited other undesirable behaviors. Although these studies were conducted in controlled environments, they sparked discussions about how advanced AI systems might pursue the goals set for them.

“Here lies one of the biggest challenges in developing increasingly powerful AI. The system does not necessarily do what we intended – it does what allows it to achieve the programmed result. Therefore, the more complex the system and the more autonomous actions we allow it to perform, the more important not only its capabilities but also control becomes,” says the expert.

AI learns from humans, and they are not neutral

The better the AI, the more it can do, but at the same time, its ability to find solutions that humans did not anticipate increases. Modern AI not only generates text but can also perform actions and use various tools. Therefore, an inaccurate answer and a wrong system action are risks of different nature.

Read more Minister of Energy: compensation of up to 200 euros must be allocated to residents left without electricity

“If a chatbot provides an incorrect answer, that is one problem. If we give an AI agent access to email, documents, company systems, or other tools and it misinterprets its task, the consequences can be entirely different. Therefore, AI development is not just about how smart it will become. Equally important is how many actions we entrust it to perform autonomously,” says A. Bagdonavičienė.

Another important aspect is that AI learns from human-created content. And humans are not neutral: the internet is full of errors, stereotypes, manipulations, aggressive content, and various biases. This means that AI is not “born” with our values. It has to, in a sense, replicate them from data, the training process, and later human-determined behavior.

Old security principles may no longer suffice

However, the expert urges not to get too involved in blind technological optimism or catastrophic scenarios. Current experiments with manipulative or rule-bypassing AI behavior are often conducted in specially created, controlled environments. This is not proof that AI wants to harm humans or that a scenario is approaching where it becomes an autonomous villain. Still, it is an important signal that as AI capabilities grow, old security principles may no longer be enough.

At the same time, it should not be forgotten that AI is only one part of the modern cybersecurity equation. In practice, organizations and residents more often face not theoretical AI threats but fraud attacks, malicious links, data theft, or other everyday digital dangers.

According to the interviewee, as AI capabilities increase, it is important not only to ensure that the systems themselves operate reliably but also to take care of broader digital security. Therefore, it is important to use additional protective measures such as Telia Safe, which helps block dangerous websites, provides extra protection when connecting to banks, warns about possible personal data leaks online, and helps browse more safely on both phones and computers. However, the challenges posed by technology are not limited to everyday cyber threats.

“We should not say that AI is becoming ‘bad.’ It would be more accurate to say that it is becoming more capable, and at the same time, our mistakes in defining its goals can become more costly. While the system only generates text, the risk is one thing. When it can autonomously make a sequence of decisions and perform real actions, the question ‘did it understand what we wanted from it?’ becomes critically important,” says the Telia expert.

Read more «Top» league – extreme draw in Šiauliai and important victory for Vilnius team

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *