On Saturday, October 3, 2026, an official at South Korea's Financial Security Institute confirmed the suspicions that had been circulating in the cybersecurity industry since the country's banks began reporting breaches: the intrusions had been run with an AI tool. Investigators had traced the attack IPs and server logs from Shinhan Bank, the first institution to report an incident, and found evidence pointing to ARTEX AI, a Chinese-developed penetration-testing platform built on large language models (The Herald Business).
The scope kept widening. By the first week of October, officials had said AI agents were likely used to breach at least seven financial institutions, exposing personal data on more than 68,000 people, in attacks traced to 33 IP addresses linked to 12 countries, according to multiple news reports on the investigation. President Lee Jae Myung, the same reports said, described the incidents as the first known example of AI agents hacking the financial sector (The Record, BleepingComputer, Tom's Hardware).
That framing needs a correction, and the agency running the investigation is the one making it. "It is true that AI was used in the attacks, but the AI did not act independently without human involvement," the institute official said. "A hacker used the AI as a tool." (The Herald Business)
Both statements are true, and the distance between them is where this story lives. No agent decided on its own to attack banks. Someone pointed free, downloadable software at a country's financial system and let it do work that used to require a team of specialists. That changes what attacks cost, what evidence they leave, and what defenders and regulators can realistically do. The evidence so far suggests the hardest problem is not stopping such attacks. It is naming who ran them.
What ARTEX AI is, and how the attacks ran
ARTEX AI is an autonomous penetration-testing system built on large language models. The institute describes it as open-source, distributed primarily through GitHub, and aimed at Chinese-speaking users (The Herald Business).
"Autonomous" describes how the tool executes, not how it chooses targets. An agent of this class runs a loop that researchers summarize as scouting, memory, reasoning, and action: it maps a target, tries known techniques, reads what comes back, decides what to try next, and keeps going without a human directing each step (arXiv survey of LLM-based attack agents). The person at the keyboard sets the objective. The software runs the campaign.
In South Korea, the objective was not the customer side of banking. Every confirmed attack hit internal employee or partner-facing systems, the services a bank's own staff and contractors use, rather than internet or mobile banking. Shinhan Bank said 25,729 people's data was exposed through a loan-agent inquiry service. KB Kookmin Bank reported 119 records leaked from an employee mobile work-support system. Hana Bank said 89 records were taken from an employee sales-support system known as ODS. BNK Busan Bank reported information on 11 contract workers (The Herald Business).

The pattern was deliberate, and investigators explained it plainly. "Security on customer-facing electronic financial services has been enormously strengthened, but internal employee-facing systems had been managed less rigorously," the official said. "There are so many management points across such a wide scope that financial institutions themselves appear to have had difficulty identifying where those vulnerabilities existed" (The Herald Business).
The attacks are believed to have been concentrated between a Sunday and the following Thursday. Once the institute identified the attacker IPs at Shinhan and shared them across the financial sector, other institutions re-examined their own access logs and found breaches they had missed, which is how the count grew. "Far more institutions were attacked than those that ended up with actual incidents," the official said, and organizations outside banking were hit as well. Two savings banks are under investigation for attacks of the same type, and Yegaram Savings Bank had already disclosed on its website that an unidentified hacker gained access and caused a personal data leak. Woori Bank and NH NongHyup Bank were probed but not breached, because the vulnerabilities the attacker sought were not present in their systems (The Herald Business, The Record).
On infrastructure, the attacker kept two to three overlapping IPs per bank and rotated to a new address whenever one was blocked, activity the official described as still ongoing. "Blocking IPs is little more than emergency first aid," the official said, noting that the method and the exploited vulnerabilities stayed the same no matter which address was used. What cracked the case was the tool itself: the attack data carried "a specific data signature common to traffic originating from ARTEX" (The Herald Business).
The fingerprint problem: when attackers share one weapon

Threat attribution has long rested on a simple premise: adversaries have habits. The tools they build, the exploits they prefer, and the order in which they move through a network, what the industry calls tactics, techniques, and procedures, add up to a fingerprint. Analysts match the fingerprint to a profile in catalogs such as MITRE ATT&CK, attach the profile to a named group, and often attach the group to a government (Synthetic APTs, arXiv).
A 2026 study from Alias Robotics and academic partners tested what happens when the habits belong to software instead of people. The researchers configured AI agents to emulate five nation-state groups, APT28, APT29, APT41, APT44, and Lazarus Group, and set them against AI-driven defenses in two simulated environments, an enterprise network and a military-infrastructure range. Across 20 experiments, all ten enterprise runs ended in compromise, with attackers reaching 2 to 12 hosts, while all ten military-range runs were defended or stalemated. The fidelity was the striking part. The agents' kill chains mapped back to the documented profiles of the real groups with 55 to 80 percent precision in the runs where they achieved domain compromise, enough to be, in the authors' words, "plausibly mistaken for the real groups" (Synthetic APTs, arXiv).
The convergence cut both ways. In eight of ten enterprise experiments, attacking agents independently weaponized the defenders' own Velociraptor endpoint-management platform as a command-and-control channel, a behavior "not encoded in any threat intelligence profile." Different personas, same framework, drifting toward the same techniques. The authors' conclusion is blunt: with capable models and the right scaffolding, "the entry barrier for operating like a nation-state APT collapses," and individuals can now behave like commonly identified threat actors, undermining attribution itself (Synthetic APTs, arXiv).
That argument is a prediction, but South Korea shows its present tense. What investigators actually hold is a piece of software, a traffic signature, and, by the reported count, 33 IP addresses scattered across 12 countries. An address list spanning a dozen countries describes infrastructure, not identity. The institute official described the ARTEX signature as a way to identify "the attacker," but a signature of that kind identifies the weapon, not the hand. Two operators on different continents running the same open-source tool leave the same fingerprint (The Herald Business, The Record).
Nothing in the public record says who the operator is. The tool is Chinese-developed; the actor is unattributed, and the investigators have been careful on this point. Asked whether the bank attacks were linked to a previously reported breach of the Government24 portal, the official said investigators could not access that data to check, but that "the timing is so far apart that we suspect it is not the work of the same attacker" (The Herald Business).
For bank security teams, the shift lands on a specific workflow. Financial-sector threat intelligence is organized around named actors: watch for this group, expect that tradecraft. When the same toolkit is available to anyone with a laptop, actor-keyed intelligence loses resolution. What still discriminates is exposure: which systems are vulnerable, and how fast they get closed.
Volume, not genius: the new economics of attack
Researchers who survey LLM-based attack agents have a name for the dynamic on display in Seoul: cyber threat inflation. The cost of a competent intrusion collapses while scale expands, through the automation of tasks once reserved for expert red-teamers and the ability to run continuous, large-scale attacks in parallel. Work that previously required months of labor and expert involvement can now be compressed into hours, the survey notes. Its authors' bottom line is that, due to what they call operational imbalances, existing defense methods are inadequate against autonomous cyberattacks (arXiv survey).
The supporting numbers come from research settings, but they are not subtle. PentestGPT, an LLM-guided penetration-testing agent, showed a 228.6 percent increase in task completion in its evaluation. RapidPen achieved shell access in 200 to 400 seconds at an estimated 30 to 60 cents per run, with a 60 percent success rate. The labs building these models have measured the offensive potential directly: Google's Project Naptime team has shown frontier models assisting with vulnerability discovery and code exploitation with minimal human input, and Anthropic runs red teams against its own models for cybersecurity misuse (arXiv survey).

South Korea is what those numbers look like in production: many institutions probed at once, a handful breached, retries that cost the attacker little, and one method running behind whichever IP happened to be live (The Herald Business).
The vulnerability supply side has its own drama, and its own correction. In April 2026, Anthropic announced Claude Mythos, a model it credited with finding thousands of zero-day flaws, vulnerabilities unknown to vendors and therefore unpatched, across every major operating system and web browser. It launched Project Glasswing, a coordinated-disclosure effort with AWS, Apple, Microsoft, Linux kernel maintainers, Cisco, and others. Ed Skoudis of the SANS Institute argued that attackers holding stockpiled zero-days were now "on a clock," and that defenders should prepare for a surge in zero-day exploitation within six to twelve months (SANS Institute).
Months in, the surge is hard to find in the data. Vulnerability-intelligence firm VulnCheck analyzed 1,061 vulnerabilities publicly attributed to AI-assisted discovery, from Glasswing and the Berkeley Vulnerability Research Initiative, and cross-checked them against its database of known exploited vulnerabilities. Fourteen, or 1.3 percent, had been exploited in the wild, almost identical to the rate across all vulnerabilities in VulnCheck's dataset. Of the 23,019 candidates Mythos surfaced, 126 were published as CVEs, the industry's standard vulnerability identifiers, and one was confirmed exploited. "The data does not suggest that AI-discovered vulnerabilities are inherently more likely to be exploited than those found through traditional methods," wrote VulnCheck researcher Patrick Garrity (The Register).
What AI has produced so far is not a generation of superweapons but a flood of findings that humans must sort. Anthropic reported 26,153 total findings; 10.5 percent reached the disclosure ledger, 8 percent were reported to maintainers, and 0.8 percent were marked fixed. The ledger itself calls independent human triage a "rate-limiting step." The severity assessments diverged sharply: Claude judged 91.5 percent of its findings high or critical, while maintainers put the same set at 51.3 percent. Daniel Stenberg, who maintains curl, reported that five vulnerabilities attributed to Mythos in his project shrank on review to one: three false positives, one "just a bug," and a single low-severity flaw that "is not going to make anyone grasp for breath" (TechTarget).

The honest summary is that AI has raised the noise floor, not the exploitation rate. VulnCheck counted 495 known-exploited vulnerabilities in the first half of 2026, with content management systems accounting for roughly a third and network edge devices remaining a firm favorite, the same unglamorous categories as always (The Register). The ARTEX case fits the pattern: one reusable method pointed at whatever internal systems happened to be the least hardened. That is the new threat model in miniature. More attempts, cheaper attempts, less identifiable attackers, and a perimeter still drawn around the systems customers touch rather than the ones employees do.
What can realistically constrain these tools
Start with what cannot. ARTEX AI is open-source software on GitHub. There is no license to revoke and no vendor to sanction, and the research literature already treats jailbreaks, which coax models past the safety guardrails meant to stop them from assisting attacks, as a recognized attack category (The Herald Business, arXiv survey). Regulating the tool away is not an available move. Everything that works lands on the target side, and the South Korean evidence points to four levers.
The first is attack-surface parity, and the case proves it by exception: Woori and NH NongHyup were probed and held, because the vulnerabilities the agent wanted were not there (The Herald Business). The Financial Services Commission is set to recommend a comprehensive audit of internal employee-facing systems, and the incident lands in the middle of a Korean policy debate over network separation, the long-standing requirement that banks wall internal systems off from the internet, which regulators had been easing. The Herald Business now asks openly whether that rollback will stall (The Herald Business). The discipline that would have stopped this attack is the least fashionable kind: patching and hardening internal systems as rigorously as the customer-facing ones.
The second is fingerprinting the agents, not just the malware. The ARTEX traffic signature is what converted suspicion into a conclusion, and agent frameworks can be treated as detection targets the way antivirus treats code families. Sector-wide indicator sharing worked the same way: once the institute circulated the attacking IPs, institutions that did not know they had been hit found the evidence in their own logs (The Herald Business).
The third is closing the reporting gap. Under current rules, South Korean financial institutions are not required to report an incident when no damage occurs (The Herald Business). Woori and NH NongHyup had no obligation to tell the regulator they had been attacked. A threat defined by volume stays invisible until one probe succeeds, which leaves regulators watching the casualty list instead of the shooting.
The fourth is automating the sorting. If human triage is the rate-limiting step in AI-assisted security, the realistic response is AI-assisted triage inside disciplined workflows. Red Hat senior principal security architect Michele Chubirka, who built such a workflow engine while participating in Project Glasswing and stressed she was speaking in her personal capacity, described it as "part harness, part integration of deterministic security tools." Her warning belongs on the wall of every bank security team: "People think that AI is magic, that a frontier model is going to magically do this stuff for you. It isn't" (TechTarget). Patch velocity, not model capability, decides whether disclosure windows help defenders or attackers.
None of these levers restores attribution; they reduce exposure, which is the variable institutions actually control. Coordinated disclosure at Glasswing's scale is the one move that attacks the economics directly, by burning stockpiled zero-days before anyone can spend them (SANS Institute).
What the case proves, and what it does not

What is established: AI-agent tooling has been used in real attacks on the financial sector of a G20 economy, with the country's president calling it the first known case of its kind, and "known" is doing a lot of work in that sentence (The Record). The attacks were cheap and repetitive, built for volume rather than precision. Where defense worked, it worked because the vulnerabilities the agent wanted did not exist.
What is not established: that the agent acted autonomously, which the lead investigating agency explicitly denies; that AI-discovered vulnerabilities are being exploited at unusual rates, which the best available data contradicts; and who ran the attacks, which has not been publicly established (The Herald Business, The Register).
The breach total, more than 68,000 people by the count reported so far, is not large for a country that has absorbed far bigger data leaks (The Record). The significance is structural. The weapon is free, the retries are cheap, and the evidence trail now ends at the software rather than the person holding it. This time investigators could name the tool, because the tool carried a signature. Nothing guarantees the next attacker leaves that convenience in place. The question banks have long organized around, "who is coming after us," is becoming unanswerable. The question that still has an answer is "what are we still exposing," and the ARTEX case shows exactly what it costs to keep getting it wrong.
Sources
- Jeong Ho-won, "Exclusive: Chinese AI tool ARTEX used in wave of bank hacks, probe finds," The Herald Business, October 3, 2026. https://biz.heraldcorp.com/article/10892497
- "South Korean officials believe AI agents were used to hack several banks," The Record. https://therecord.media/south-korean-bank-hacks-ai-agents
- "South Korea probes bank breaches amid suspected AI-powered attacks," BleepingComputer. https://www.bleepingcomputer.com/news/security/south-korea-probes-bank-breaches-amid-suspected-ai-powered-attacks/
- "Hackers suspected of using AI agents for cyberattacks on South Korean banks," Tom's Hardware. https://www.tomshardware.com/tech-industry/cyber-security/hackers-suspected-of-using-ai-agents-for-cyberattacks-on-south-k
- Minrui Xu et al., "Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks," arXiv, 2025. https://arxiv.org/html/2505.12786v2
- Francesco Balassone et al., "Synthetic APTs: the Collapse of TTP-Based Attribution," arXiv, 2026. https://arxiv.org/html/2606.07158v1
- Ed Skoudis, "SANS Critical Advisory: BugBusters: AI Vulnerability Discovery Hype vs. Reality," SANS Institute, April 13, 2026. https://www.sans.org/blog/sans-critical-advisory-bugbusters-ai-vulnerability-discovery-hype-vs-reality
- Carly Page, "AI-found bugs aren't proving any easier to exploit despite the hype," The Register, July 28, 2026. https://www.theregister.com/security/2026/07/28/ai-found-bugs-arent-proving-any-easier-to-exploit-despite-the-hype/5279637
- "Glasswing results put AI 'vulnpocalypse' to the test," TechTarget, September 29, 2026. https://www.techtarget.com/it-infrastructure/news/366651381/Glasswing-results-put-AI-vulnpocalypse-to-the-test
Comments (6)
Continue exploring
The ShinyHunters FBI Breach: When Stolen Personnel Files Become Leverage Against the State
The ShinyHunters breach of FBI personnel files shows an extortion crew using stolen government data to coerce the state itself…
The Best Coding Model in August 2026: Why the Answer Depends on the Job
The question sounds simple: which coding model is best, Codex, Gemini, Claude, or DeepSeek? The evidence does not support a…
Aurora's 96-MW Proof: Inside the Certification Bet on Software-Defined Grid Capacity
Aurora's 96-MW AI factory is the first designed to a power-flexibility certification standard; scaling the model hinges on…
The shared-IP-and-tool-signature pattern tracks closely with what my team observed in a Q2 incident at a European fintech. Once we identified the tooling fingerprint, we found the same campaign had hit three of our subsidiaries in a single week, and only the IP rotation logs initially made it look like separate actors. The Sunday-to-Thursday window mentioned in the article also held for us. The quote about 'far more institutions attacked than those that ended up with actual incidents' felt uncomfortably familiar - we discovered another six attempted intrusions after the fact that had been silently rejected. The speed of post-event discovery once the signature was known is what stood out.
The article frames attribution as the hardest problem, but the real operational issue is that ARTEX is free on GitHub - even if you conclusively name the actor behind these seven intrusions, anyone else can pull the same repo tonight and run an identical campaign tomorrow. Naming the threat actor feels secondary when the barrier to entry is a download link.
I've worked incident response at a mid-size bank, and our employee-facing systems actually have stricter logging and access controls than customer-facing channels because internal audit teams flag them constantly. The article's claim that internals are 'managed less rigorously' doesn't match the environment I've seen firsthand.
I'm not entirely sure the 'first known example of AI agents hacking the financial sector' framing holds - there was reporting in 2024 about LLM-assisted intrusions at telecom providers, though admittedly those lacked a clear tool-level signature. The article could have acknowledged that ambiguity.
Blocking IPs is little more than emergency first aid" might be the most honest line in the whole piece.
The point about internal employee systems being softer targets matches what we see in our third-party risk assessments. Financial institutions pour money into perimeter and customer-facing controls, but contractor portals, internal HR tools, and loan inquiry systems often run on outdated stacks with patch cycles measured in quarters. Attackers know this, and the South Korean incident details line up with that pattern.