#ai #future #cyber
Just a fortnight back, OpenAI’s model in development broke out of isolated test environments. These test conditions were previously unknown.
Now Anthropic’s three models, including Mythos 5, broke into random systems. As per company this has happened starting since April, but now they are publishing it.
The Anthropic models involved in this cyber attack are:
I. Mythos 5,
II. Opus 4.7, and
III. An ongoing research test model.
Mythos 5 and Claude Fable 5 are the exact same underlying AI model, but they are configured with radically different safety guardrails. Yes safety guardrails.
The test environment was told the internet is not available. But the internet was available. They started cyber access. So whose fault is it then?
Is this negligence?
These super-intelligent Anthropic models could not have been fooled. They started cyber access.
The aim was a cyber challenge where the models were told to find a “flag” hidden in a different machine. And the system was told to find the flag with no restrictions.
We all know how intelligent these models are! They have replaced so many jobs!
The models learned the internet was available; still, it did not stop!
Anthropic restated that these situations may have been averted by steps taken by them and their third-party partners.
Even OpenAI’s development model found out about the internet when the internet was not readily available; here, the internet was clearly there.
This could have been prevented.
Luckily, it did not cause as much harm as it could have, given Mythos 5 is one of the most powerful models in the world!
What if it was out of a test environment?
Prior Mythos 5 has been declared safe by the company!
Are we safe using these models? Given they are autonomous?
Should we continue our use of such models?
We hope the companies would make models safer.
Coming Soon. How Anthropic’s attacks differ from OpenAI’s!