SOCFRAME / THRONE Journal
All reports

Scenario 07 · Live operation · Nation-state espionage

We stopped scripting.
We drove a live adversary.

For six scenarios we executed techniques by hand. This time we handed the wheel to a real command-and-control server and let it fight. Caldera planned and chained an autonomous APT29 operation off the facts it discovered — credential theft, discovery, defense evasion, lateral movement — and THRONE caught the whole kill chain as it happened, in real time.

Adversary
APT29 · Cozy Bear
Mode
Autonomous · fact-chained
Operation
throne-apt29-emu
Result
Detected end to end
Executive summary

An autonomous nation-state intrusion, caught end to end

APT29 — the Russian SVR intrusion set MITRE tracks as G0016 and the public knows as Cozy Bear — is one of the most capable espionage actors in the world. For this scenario we did not script its techniques; we ran MITRE's own APT29 emulation plan as a live, autonomous operation and let the planner decide every move.

A Caldera v5.3.0 command-and-control server drove an elevated agent on a Windows victim with no human in the loop: the planner read the facts the agent collected and chained the next applicable abilities itself. THRONE observed the whole operation through Sysmon telemetry, matched it against its Sigma and behavioral engines, and ONYX — THRONE's AI SOC analyst — auto-triaged the resulting alerts into scored, grouped incidents, each one rebuilt as a process-causality tree.

The headline: the complete kill chain was detected as it ran. Credential access, discovery, execution, defense evasion and attempted lateral movement all raised alerts within seconds of execution. The single most important catch, INC-55853, was a CRITICAL behavioral detection of LSASS memory access (T1003.001) — in-memory credential theft identified by behavior, not by a signature an attacker can simply rename around.

Headline result

The operation ran fully autonomously from the C2 — 85 planned decisions, 0 human steps — and THRONE produced scored, investigable incidents for every detected tactic. The crown-jewel moment, an attacker reaching into LSASS for every credential on the box, was flagged CRITICAL at 87% confidence the instant it happened.

The adversary

APT29 — Cozy Bear, the SVR's quiet professionals

APT29 (G0016) is attributed to Russia's Foreign Intelligence Service, the SVR. It is a long-running espionage program, not a smash-and-grab crew — patient, stealthy, and relentlessly focused on governments, diplomatic missions, think tanks and the technology supply chain.

Attribution and aliases

The group is tracked across the industry under many names: Cozy Bear, The Dukes, YTTRIUM, UNC2452, Dark Halo, and — in Microsoft's taxonomy — Midnight Blizzard (formerly Nobelium). US and UK governments have publicly attributed its operations to the SVR. Its defining trait is restraint: APT29 generally prefers to look like a legitimate user than to make noise, which is exactly what makes it a hard detection problem.

Notable real-world campaigns

The record is long and consequential:

• The Dukes (from roughly 2008): a family of toolkits — MiniDuke, CosmicDuke, CozyDuke, SeaDuke, HAMMERTOSS — used in sustained campaigns against Western governments and NGOs.
• 2016 DNC compromise: APT29 was one of two Russian groups operating inside the Democratic National Committee.
• 2020 SolarWinds / SUNBURST: the defining APT29 operation — a supply-chain compromise (T1195.002) that trojanized SolarWinds Orion updates (SUNBURST, then TEARDROP), reached thousands of organizations and multiple US federal agencies, and pivoted into cloud identity, including forged SAML tokens / Golden SAML (T1606.002).
• 2020 vaccine research: targeting of COVID-19 vaccine developers with the WellMess and WellMail implants (joint NCSC / NSA / CSE advisory).
• 2021 USAID: spearphishing at scale through a hijacked Constant Contact email-marketing account.
• 2023–2024 Midnight Blizzard: password-spray (T1110.003) into a legacy, non-production tenant, then OAuth application abuse (T1098) to read corporate email at Microsoft and other targets.

Signature tradecraft

Across those campaigns the same playbook recurs: initial access by spearphishing or supply chain; a strong preference for living off the land and valid accounts (T1078) over noisy custom malware; an aggressive focus on the identity and cloud plane — stealing tokens, abusing OAuth application consent and service-principal credentials, password spraying, and forging SAML assertions; and meticulous defense evasion — hiding payloads in alternate data streams, compiling code on the victim after delivery, timestomping, and clearing indicators. When APT29 does reach for credentials on an endpoint, LSASS memory (T1003.001) is a favoured source. It is an adversary that wins by blending in, not by overwhelming.

What we actually emulated

MITRE's Caldera emu plugin ships an APT29 profile derived from the ATT&CK Evaluations (Round 2) emulation plan — a two-part script modelled on the Dukes: a rapid "smash-and-grab" day of collection and exfiltration, and a stealthy "low-and-slow" day of persistence and lateral movement. That is the plan we drove autonomously here, letting the planner, not an operator, decide the order.

Setup & methodology

A real C2, not a shell script

The previous scenarios proved THRONE sees individual techniques. The harder question: can it follow an adversary that decides its own next move? So we drove the real thing — autonomously, fact-chained, with no human issuing commands.

The attacker side

A Caldera v5.3.0 command-and-control server at 178.104.87.27 running the MITRE emu emulation plugin. We launched its APT29 profile — MITRE's ATT&CK-Evaluations Cozy Bear plan — with the batch planner in autonomous mode and the APT29 (Emu) fact source, as operation throne-apt29-emu. "Autonomous" is literal: the planner reads the facts the agent collects — hostnames, usernames, domain details, file paths, tickets — and selects and chains the next applicable abilities itself. No operator typed a command after launch. Across the run it planned 85 decisions.

The target side

On the victim, a Caldera Sandcat agent (hkiejm, elevated) beaconing over HTTP from the Windows host WIN-2V8LL1URVNV (2.29.105.165). The host runs Sysmon and streams its telemetry, tenant-tagged, to the THRONE cloud tenant. This is a lab environment, and everything below — every IP, hostname, agent name and incident ID — is shown exactly as captured. Nothing is masked.

85
Decisions planned
25+
Abilities succeeded
1
Elevated agent
0
Human steps

The detection pipeline

Every ability the planner fired on the victim generated Windows telemetry, and that telemetry travelled the same automatic path on every event:

Caldera planner → ability on WIN-2V8LL1URVNV → Sysmon event
→ tenant-tagged syslog → THRONE :1514 → Kafka → ch_pump parse → ClickHouse
→ Sigma sweep (3,148 rules / 389 ATT&CK techniques) + behavioral engine → alert
→ ONYX auto-triage & grouping → incident → causality tree (parent/child process GUIDs)

No analyst touched any of it until there was a scored, grouped incident with a process tree waiting. That is the loop THRONE is built to close: adversary behaviour in, an investigable incident out.

An honest note on the lab

Because this is a single-victim lab, the lateral-movement (RDP / SMB) and exfiltration (OneDrive) abilities fired but had no second host or cloud credential to land on — expected, and marked "attempted" below. The value is the telemetry those abilities threw and whether THRONE saw it. It did. Launch was at 15:48Z on 2026-10-09.

The kill chain

What the autonomous planner chained — phase by phase

A sample of the abilities Caldera executed on WIN-2V8LL1URVNV, by tactic. Each generated Windows telemetry that THRONE matched against its Sigma and behavioral detections. The order was not ours — the planner chose it off the facts it discovered.

Detected
Credential accessInvoke-Mimikatz
Detected
Credential accessLSASS memory read
Detected
Discoverywhoami · net · nltest
Detected
Discoverysysteminfo · klist
Detected
ExecutionPowerShell · WMIC
Detected
Defense evasionCsc.EXE .NET compile
Detected
ExecutionWinAPI via cmdline
Detected
Lateral movementMstsc RDP · SMB
Attempted
PersistenceStartup folder
Attempted
ExfiltrationOneDrive

Execution — PowerShell, WMI and the Native API

The operation ran its tradecraft through trusted interpreters: PowerShell (T1059.001), the Windows command shell (T1059.003), WMI / WMIC (T1047), and the Windows Native API invoked straight from a command line (T1106). WMI spawns and queries processes without a conventional CreateProcess; Native API calls such as VirtualAlloc and CreateThread, reached via Add-Type / [DllImport], execute shellcode in memory. Why APT29 uses it: living off the land. Every one of these is a legitimate administrative primitive, so the activity blends into normal operations and leaves little on disk to scan.

Credential access — the LSASS prize

The run reached for credentials two ways: Invoke-Mimikatz and a direct read of LSASS process memory (T1003.001), plus SAM / ticket material (T1003.002) and Kerberos ticket enumeration with klist. Reading lsass.exe memory lifts plaintext passwords, NTLM hashes and Kerberos tickets in one move; klist lists cached tickets to find credentials worth reusing. Why APT29 uses it: harvested credentials feed valid-account access (T1078) — the quiet lateral movement they prefer over exploits. This is the step the whole intrusion is built around, which is why it is the one THRONE weights most heavily.

Discovery — mapping the ground for the planner

A full reconnaissance sweep ran with built-in tools: whoami (T1033), net / net1 (T1087 accounts, T1069 groups, T1016 network config), nltest (T1482 domain-trust discovery), systeminfo (T1082) and klist. Together they enumerate the current user, local and domain accounts and groups, domain trusts, OS build and cached tickets. Why it matters here specifically: these facts are exactly what the autonomous planner consumes to decide its next move — in this operation, discovery is both tradecraft and the engine driving everything else. APT29 keeps to native tools so the recon reads like routine administration.

Defense evasion — compile on the box, hide in the stream

Two evasion techniques stood out: dynamic .NET compilation with Csc.EXE (T1027.004, Compile After Delivery) and PowerShell executed from an NTFS alternate data stream (T1564.004). The first ships source code rather than a binary and compiles it on the victim with the trusted .NET compiler, so there is no pre-built malicious file for static scanning to catch; the second stashes scripts in an ADS, where ordinary directory listings don't show them. Why APT29 uses it: minimise signature-based detection and forensic artifacts — consistent with a group that values staying unseen over moving fast.

Lateral movement — stolen creds, trusted protocols

The planner reached for Mstsc.EXE RDP (T1021.001) and SMB / administrative shares (T1021.002). Why APT29 uses it: move with harvested valid accounts over legitimate remote protocols rather than exploits — far quieter than dropping tooling on each new host. In this single-victim lab there was no second machine to reach, so these fired as "attempted" — but the telemetry still hit THRONE.

Persistence and exfiltration

Finally, startup-folder persistence (T1547.001) and exfiltration to OneDrive (T1567.002). The first drops a launch item that runs at logon; the second stages stolen data out through a sanctioned cloud service. Why APT29 uses it: survive reboots unobtrusively and blend exfiltration into normal cloud traffic — the identity-and-cloud centre of gravity that runs through every major Dukes campaign. Both fired as "attempted": no persistent foothold worth keeping and no cloud credential to land on in the lab.

The detections

What THRONE caught — in real time

Within seconds of each ability running, Sigma and the behavioral engine fired, and ONYX — THRONE's AI SOC analyst — auto-triaged each alert into an incident with a confidence score. The wave from this one operation, all on WIN-2V8LL1URVNV, timestamped to the launch at 15:48Z:

IncidentDetectionEngineSevATT&CKONYX
INC-55853[IOA] LSASS Memory Accessbehavior_engineCRITT1003.00187%
INC-55847[Sigma] Mimikatz Usesigma_engineHIGHT1003.00287%
INC-55846[Sigma] Suspicious Program Namessigma_engineHIGHT1059.00187%
INC-55849[Sigma] Potential WinAPI Calls Via CommandLinesigma_engineHIGHT110672%
INC-55851Multiple alerts — defense-evasionsigma_engineHIGHTA0005—
INC-55850Multiple alerts — discovery (3)sigma_engineMEDTA0007—
INC-55852Multiple alerts — execution (2)sigma_engineMEDTA0002—
INC-55848Multiple alerts — defense-evasionsigma_engineMEDTA0005—

The alert stream behind those incidents named the exact tradecraft: Csc.EXE dynamic .NET compilation, Net.EXE group/account recon, systeminfo, Mstsc.EXE RDP, PowerShell run from an alternate data stream, and WMI process creation — the APT29 playbook, caught step by step.

INC-55853 — LSASS memory access, caught by behavior (CRITICAL)

The behavioral engine does not wait for a named tool. It keys on the action: a non-system process obtaining a read handle into the memory of lsass.exe (surfaced by Sysmon's ProcessAccess event), which is the technical heart of credential dumping. That maps to T1003.001 (OS Credential Dumping: LSASS Memory). Because the signal is the memory access itself, it fires whether the tool is Mimikatz, a renamed copy, a comsvcs.dll MiniDump, or bespoke code — exactly the variability APT29 brings. ONYX triaged it into a CRITICAL incident at 87% confidence, the highest-severity event of the operation and, operationally, the one that should pull an analyst first.

INC-55847 — Mimikatz use, with a rebuilt process tree (HIGH)

The Sigma rule "Mimikatz Use" keys on the command-line and image telltales of Mimikatz-family tooling and Kerberos interaction; here it fired on powershell.exe spawning whoami and klist to enumerate cached Kerberos tickets. It mapped to T1003.002, was raised as alert ALR-175473 at 87% confidence, and was grouped by ONYX into INC-55847 — the same incident whose causality tree appears in the evidence below.

INC-55849 — Native API calls via command line (HIGH)

The Sigma rule "Potential WinAPI Calls Via CommandLine" keys on PowerShell invoking Win32 APIs directly — Add-Type / [DllImport], GetProcAddress, VirtualAlloc, CreateThread — the fingerprint of in-memory execution that avoids spawning new processes. It mapped to T1106 (Native API). ONYX scored it 72%: a strong signal, but on its own lower-confidence than a credential-theft event, and the triage reflects that calibration rather than treating every HIGH alert as equal.

INC-55846 — suspicious program names (HIGH)

The Sigma rule "Suspicious Program Names" keys on process images whose names match known-malicious or deliberately misleading patterns the emulation introduced. It mapped to T1059.001 and ONYX triaged it HIGH at 87%. It is a useful reminder that signatures still pull their weight on the noisy, tool-driven parts of an intrusion — while the behavioral engine covers the parts a signature can't anticipate.

INC-55850 / INC-55852 / INC-55848 — clustered into tactic-level incidents

Not every detection deserves its own incident. ONYX grouped the discovery alerts (whoami, net / net1, systeminfo, klist) into INC-55850 (TA0007, MED), the execution pair into INC-55852 (TA0002, MED), and the defense-evasion alerts into INC-55848 and INC-55851 (TA0005). Grouping is the difference between an analyst facing eight disconnected alerts and facing a handful of incidents that each tell one coherent story.

The one that matters most

INC-55853 is the crown jewel: a CRITICAL behavioral detection of LSASS memory access (T1003.001) — the in-memory credential theft at the heart of Mimikatz, caught by behavior, not just a signature. This is the moment an intruder reaches for every password on the box, and THRONE flagged it the instant it happened.

The evidence

Every incident, rebuilt as a process tree

THRONE reconstructs each incident's causality from Sysmon's process GUIDs — every process-creation event carries a ProcessGuid and a ParentProcessGuid, and THRONE stitches them parent-to-child into the full lineage, then renders it live on a 2D canvas. No log-grepping; the analyst sees the whole tree. Captured directly from the Incidents tab during the operation.

THRONE causality graph for INC-55850 showing powershell.exe and cmd.exe spawning systeminfo, net, net1, whoami and klist
INC-55850 — Discovery. A 9-node causality tree: powershell.exe (PID 3936) and two cmd.exe processes fan out to systeminfo.exe, net.exe → net1.exe, whoami.exe and klist.exe (×2) — the full Cozy Bear reconnaissance sweep, linked into one investigable incident.
THRONE causality graph for INC-55847 showing powershell.exe spawning whoami and two klist processes
INC-55847 — Credential access (Mimikatz Use). powershell.exe (PID 3936) spawning whoami.exe and klist.exe — Kerberos ticket enumeration flagged at 87% confidence (ALR-175473), mapped to T1003.002.

Both trees trace back to powershell.exe at PID 3936 — the same beacon-spawned shell the planner drove throughout the operation. That shared root is the point: discovery and credential access are not two unrelated alerts, they are two branches of one intrusion, and THRONE shows them that way.

Findings

Why this is the hard case

Driving a live, autonomous adversary is a different test from replaying a script. The lesson is in what made it hard — and in what THRONE needed to keep up with a target that improvised.

APT29 is hard precisely because almost every individual action is plausible on a Windows host. whoami, net, systeminfo, PowerShell, csc.exe, mstsc — administrators run all of these every day. The malice is in the sequence and the causality: a shell that reads LSASS, then enumerates tickets, then probes domain trusts, then reaches for RDP. A detection strategy built only on atomic signatures raises a scatter of low-confidence alerts and misses the narrative that ties them together.

Three things carried the detection here. First, behavior over signature: INC-55853 caught LSASS access by what it did, so a renamed or custom credential tool would not have slipped past. Second, causality: rebuilding process trees from Sysmon GUIDs turned a wave of atomic alerts into a few investigable incidents with a common root. Third, autonomous triage: ONYX scored and grouped every alert the instant it landed, so the first thing a human would see is a ranked set of incidents — CRITICAL credential theft at the top — not an undifferentiated firehose.

2,421
Caldera abilities
41
Adversary profiles
~400
ATT&CK techniques
3,148
THRONE Sigma rules

Those four numbers describe the emulation library and THRONE's rule base — not a claim of coverage. We prove detections the only honest way: by driving real adversary behaviour at the platform and measuring what it catches, profile by profile, rather than asserting coverage on a slide. APT29 was one such profile, run end to end, measured as it happened.

Defensive takeaways

What to take back to your own SOC

Four concrete moves this operation argues for — whether or not you run THRONE.

1. Detect LSASS access by behavior, and harden the target

Alert on any non-system process opening a read handle to lsass.exe (Sysmon ProcessAccess, or your EDR's equivalent), not just on the string "mimikatz." APT29 has many ways to reach credential memory; the behavior is the constant. Then raise the cost of success: enable LSA protection (RunAsPPL) and, where you can, Credential Guard, so the memory is harder to reach in the first place.

2. Treat living-off-the-land as high-signal in context

csc.exe or MSBuild spawned by a script host, PowerShell reading from an alternate data stream, and Win32 API calls from a command line are rarely benign in sequence. Turn on PowerShell script-block and module logging plus AMSI, and score these higher when their parent is a shell or a freshly spawned beacon rather than an interactive admin session.

3. Correlate on process lineage, not just events

Individually benign discovery commands only become an intrusion when you can see they share a parent. Invest in process-tree / causality correlation — it is what collapsed eight-plus atomic alerts into a handful of incidents here, and it is what gives an analyst a story to follow instead of a queue to clear.

4. Defend the identity and cloud plane — that's APT29's real endgame

On an endpoint they steal credentials; in your tenant they abuse them. Enforce phishing-resistant MFA, audit OAuth application consents and service-principal credentials, watch for anomalous or forged tokens (the SolarWinds and Midnight Blizzard lessons), and gate RDP / SMB lateral movement behind privileged-access controls. The endpoint catch is where you stop them cheaply; the identity controls are where you stop them for good.

The throughline

Every recommendation points the same way: stop grading single events and start reading the story they tell together. That is the gap a patient, living-off-the-land adversary like APT29 is built to exploit — and the gap THRONE is built to close.