Scenario 07 · Live operation · Nation-state espionage
For six scenarios we executed techniques by hand. This time we handed the wheel to a real command-and-control server and let it fight. Caldera planned and chained an autonomous APT29 operation off the facts it discovered — credential theft, discovery, defense evasion, lateral movement — and THRONE caught the whole kill chain as it happened, in real time.
APT29 — the Russian SVR intrusion set MITRE tracks as G0016 and the public knows as Cozy Bear — is one of the most capable espionage actors in the world. For this scenario we did not script its techniques; we ran MITRE's own APT29 emulation plan as a live, autonomous operation and let the planner decide every move.
A Caldera v5.3.0 command-and-control server drove an elevated agent on a Windows victim with no human in the loop: the planner read the facts the agent collected and chained the next applicable abilities itself. THRONE observed the whole operation through Sysmon telemetry, matched it against its Sigma and behavioral engines, and ONYX — THRONE's AI SOC analyst — auto-triaged the resulting alerts into scored, grouped incidents, each one rebuilt as a process-causality tree.
The headline: the complete kill chain was detected as it ran. Credential access, discovery, execution, defense evasion and attempted lateral movement all raised alerts within seconds of execution. The single most important catch, INC-55853, was a CRITICAL behavioral detection of LSASS memory access (T1003.001) — in-memory credential theft identified by behavior, not by a signature an attacker can simply rename around.
The operation ran fully autonomously from the C2 — 85 planned decisions, 0 human steps — and THRONE produced scored, investigable incidents for every detected tactic. The crown-jewel moment, an attacker reaching into LSASS for every credential on the box, was flagged CRITICAL at 87% confidence the instant it happened.
APT29 (G0016) is attributed to Russia's Foreign Intelligence Service, the SVR. It is a long-running espionage program, not a smash-and-grab crew — patient, stealthy, and relentlessly focused on governments, diplomatic missions, think tanks and the technology supply chain.
The group is tracked across the industry under many names: Cozy Bear, The Dukes, YTTRIUM, UNC2452, Dark Halo, and — in Microsoft's taxonomy — Midnight Blizzard (formerly Nobelium). US and UK governments have publicly attributed its operations to the SVR. Its defining trait is restraint: APT29 generally prefers to look like a legitimate user than to make noise, which is exactly what makes it a hard detection problem.
The record is long and consequential:
• The Dukes (from roughly 2008): a family of toolkits — MiniDuke, CosmicDuke, CozyDuke, SeaDuke, HAMMERTOSS — used in
sustained campaigns against Western governments and NGOs.
• 2016 DNC compromise: APT29 was one of two Russian groups operating inside the Democratic National Committee.
• 2020 SolarWinds / SUNBURST: the defining APT29 operation — a supply-chain compromise
(T1195.002) that trojanized SolarWinds Orion updates (SUNBURST, then TEARDROP), reached thousands of
organizations and multiple US federal agencies, and pivoted into cloud identity, including forged SAML tokens / Golden SAML
(T1606.002).
• 2020 vaccine research: targeting of COVID-19 vaccine developers with the WellMess and WellMail implants (joint
NCSC / NSA / CSE advisory).
• 2021 USAID: spearphishing at scale through a hijacked Constant Contact email-marketing account.
• 2023–2024 Midnight Blizzard: password-spray (T1110.003) into a legacy, non-production tenant,
then OAuth application abuse (T1098) to read corporate email at Microsoft and other targets.
Across those campaigns the same playbook recurs: initial access by spearphishing or supply chain; a strong preference for living off the land and valid accounts (T1078) over noisy custom malware; an aggressive focus on the identity and cloud plane — stealing tokens, abusing OAuth application consent and service-principal credentials, password spraying, and forging SAML assertions; and meticulous defense evasion — hiding payloads in alternate data streams, compiling code on the victim after delivery, timestomping, and clearing indicators. When APT29 does reach for credentials on an endpoint, LSASS memory (T1003.001) is a favoured source. It is an adversary that wins by blending in, not by overwhelming.
MITRE's Caldera emu plugin ships an APT29 profile derived from the ATT&CK Evaluations (Round 2) emulation
plan — a two-part script modelled on the Dukes: a rapid "smash-and-grab" day of collection and exfiltration, and a stealthy
"low-and-slow" day of persistence and lateral movement. That is the plan we drove autonomously here, letting the planner,
not an operator, decide the order.
The previous scenarios proved THRONE sees individual techniques. The harder question: can it follow an adversary that decides its own next move? So we drove the real thing — autonomously, fact-chained, with no human issuing commands.
A Caldera v5.3.0 command-and-control server at 178.104.87.27 running the MITRE emu
emulation plugin. We launched its APT29 profile — MITRE's ATT&CK-Evaluations Cozy Bear plan — with the batch
planner in autonomous mode and the APT29 (Emu) fact source, as operation
throne-apt29-emu. "Autonomous" is literal: the planner reads the facts the agent collects — hostnames,
usernames, domain details, file paths, tickets — and selects and chains the next applicable abilities itself. No operator
typed a command after launch. Across the run it planned 85 decisions.
On the victim, a Caldera Sandcat agent (hkiejm, elevated) beaconing over HTTP from the Windows host
WIN-2V8LL1URVNV (2.29.105.165). The host runs Sysmon and streams its
telemetry, tenant-tagged, to the THRONE cloud tenant. This is a lab environment, and everything below — every IP, hostname,
agent name and incident ID — is shown exactly as captured. Nothing is masked.
Every ability the planner fired on the victim generated Windows telemetry, and that telemetry travelled the same automatic path on every event:
No analyst touched any of it until there was a scored, grouped incident with a process tree waiting. That is the loop THRONE is built to close: adversary behaviour in, an investigable incident out.
Because this is a single-victim lab, the lateral-movement (RDP / SMB) and exfiltration (OneDrive) abilities fired but had
no second host or cloud credential to land on — expected, and marked "attempted" below. The value is the telemetry those
abilities threw and whether THRONE saw it. It did. Launch was at 15:48Z on 2026-10-09.
A sample of the abilities Caldera executed on WIN-2V8LL1URVNV, by tactic. Each generated Windows telemetry that THRONE matched against its Sigma and behavioral detections. The order was not ours — the planner chose it off the facts it discovered.
The operation ran its tradecraft through trusted interpreters: PowerShell (T1059.001), the Windows
command shell (T1059.003), WMI / WMIC (T1047), and the Windows
Native API invoked straight from a command line (T1106). WMI spawns and queries processes without a
conventional CreateProcess; Native API calls such as VirtualAlloc and CreateThread,
reached via Add-Type / [DllImport], execute shellcode in memory. Why APT29 uses it: living off
the land. Every one of these is a legitimate administrative primitive, so the activity blends into normal operations and
leaves little on disk to scan.
The run reached for credentials two ways: Invoke-Mimikatz and a direct read of LSASS process memory
(T1003.001), plus SAM / ticket material (T1003.002) and Kerberos ticket
enumeration with klist. Reading lsass.exe memory lifts plaintext passwords, NTLM hashes and Kerberos
tickets in one move; klist lists cached tickets to find credentials worth reusing. Why APT29 uses it:
harvested credentials feed valid-account access (T1078) — the quiet lateral movement they prefer over
exploits. This is the step the whole intrusion is built around, which is why it is the one THRONE weights most heavily.
A full reconnaissance sweep ran with built-in tools: whoami (T1033), net /
net1 (T1087 accounts, T1069 groups,
T1016 network config), nltest (T1482 domain-trust discovery),
systeminfo (T1082) and klist. Together they enumerate the current user,
local and domain accounts and groups, domain trusts, OS build and cached tickets. Why it matters here specifically:
these facts are exactly what the autonomous planner consumes to decide its next move — in this operation, discovery is both
tradecraft and the engine driving everything else. APT29 keeps to native tools so the recon reads like routine
administration.
Two evasion techniques stood out: dynamic .NET compilation with Csc.EXE (T1027.004,
Compile After Delivery) and PowerShell executed from an NTFS alternate data stream (T1564.004). The
first ships source code rather than a binary and compiles it on the victim with the trusted .NET compiler, so there is no
pre-built malicious file for static scanning to catch; the second stashes scripts in an ADS, where ordinary directory
listings don't show them. Why APT29 uses it: minimise signature-based detection and forensic artifacts — consistent
with a group that values staying unseen over moving fast.
The planner reached for Mstsc.EXE RDP (T1021.001) and SMB / administrative shares
(T1021.002). Why APT29 uses it: move with harvested valid accounts over legitimate remote
protocols rather than exploits — far quieter than dropping tooling on each new host. In this single-victim lab there was no
second machine to reach, so these fired as "attempted" — but the telemetry still hit THRONE.
Finally, startup-folder persistence (T1547.001) and exfiltration to OneDrive (T1567.002). The first drops a launch item that runs at logon; the second stages stolen data out through a sanctioned cloud service. Why APT29 uses it: survive reboots unobtrusively and blend exfiltration into normal cloud traffic — the identity-and-cloud centre of gravity that runs through every major Dukes campaign. Both fired as "attempted": no persistent foothold worth keeping and no cloud credential to land on in the lab.
Within seconds of each ability running, Sigma and the behavioral engine fired, and ONYX — THRONE's
AI SOC analyst — auto-triaged each alert into an incident with a confidence score. The wave from this one operation, all on
WIN-2V8LL1URVNV, timestamped to the launch at 15:48Z:
| Incident | Detection | Engine | Sev | ATT&CK | ONYX |
|---|---|---|---|---|---|
| INC-55853 | [IOA] LSASS Memory Access | behavior_engine | CRIT | T1003.001 | 87% |
| INC-55847 | [Sigma] Mimikatz Use | sigma_engine | HIGH | T1003.002 | 87% |
| INC-55846 | [Sigma] Suspicious Program Names | sigma_engine | HIGH | T1059.001 | 87% |
| INC-55849 | [Sigma] Potential WinAPI Calls Via CommandLine | sigma_engine | HIGH | T1106 | 72% |
| INC-55851 | Multiple alerts — defense-evasion | sigma_engine | HIGH | TA0005 | — |
| INC-55850 | Multiple alerts — discovery (3) | sigma_engine | MED | TA0007 | — |
| INC-55852 | Multiple alerts — execution (2) | sigma_engine | MED | TA0002 | — |
| INC-55848 | Multiple alerts — defense-evasion | sigma_engine | MED | TA0005 | — |
The alert stream behind those incidents named the exact tradecraft: Csc.EXE dynamic .NET compilation,
Net.EXE group/account recon, systeminfo, Mstsc.EXE RDP, PowerShell run from an
alternate data stream, and WMI process creation — the APT29 playbook, caught step by step.
The behavioral engine does not wait for a named tool. It keys on the action: a non-system process obtaining a read
handle into the memory of lsass.exe (surfaced by Sysmon's ProcessAccess event), which is the technical heart of
credential dumping. That maps to T1003.001 (OS Credential Dumping: LSASS Memory). Because the signal
is the memory access itself, it fires whether the tool is Mimikatz, a renamed copy, a comsvcs.dll MiniDump, or
bespoke code — exactly the variability APT29 brings. ONYX triaged it into a CRITICAL incident at 87% confidence,
the highest-severity event of the operation and, operationally, the one that should pull an analyst first.
The Sigma rule "Mimikatz Use" keys on the command-line and image telltales of Mimikatz-family tooling and Kerberos
interaction; here it fired on powershell.exe spawning whoami and klist to enumerate
cached Kerberos tickets. It mapped to T1003.002, was raised as alert ALR-175473
at 87% confidence, and was grouped by ONYX into INC-55847 — the same incident whose causality tree appears in the
evidence below.
The Sigma rule "Potential WinAPI Calls Via CommandLine" keys on PowerShell invoking Win32 APIs directly —
Add-Type / [DllImport], GetProcAddress, VirtualAlloc,
CreateThread — the fingerprint of in-memory execution that avoids spawning new processes. It mapped to
T1106 (Native API). ONYX scored it 72%: a strong signal, but on its own lower-confidence than a
credential-theft event, and the triage reflects that calibration rather than treating every HIGH alert as equal.
The Sigma rule "Suspicious Program Names" keys on process images whose names match known-malicious or deliberately misleading patterns the emulation introduced. It mapped to T1059.001 and ONYX triaged it HIGH at 87%. It is a useful reminder that signatures still pull their weight on the noisy, tool-driven parts of an intrusion — while the behavioral engine covers the parts a signature can't anticipate.
Not every detection deserves its own incident. ONYX grouped the discovery alerts (whoami,
net / net1, systeminfo, klist) into INC-55850
(TA0007, MED), the execution pair into INC-55852 (TA0002, MED), and the
defense-evasion alerts into INC-55848 and INC-55851 (TA0005). Grouping is the difference between an
analyst facing eight disconnected alerts and facing a handful of incidents that each tell one coherent story.
INC-55853 is the crown jewel: a CRITICAL behavioral detection of LSASS memory access (T1003.001) — the in-memory credential theft at the heart of Mimikatz, caught by behavior, not just a signature. This is the moment an intruder reaches for every password on the box, and THRONE flagged it the instant it happened.
THRONE reconstructs each incident's causality from Sysmon's process GUIDs — every process-creation event carries a ProcessGuid and a ParentProcessGuid, and THRONE stitches them parent-to-child into the full lineage, then renders it live on a 2D canvas. No log-grepping; the analyst sees the whole tree. Captured directly from the Incidents tab during the operation.
powershell.exe (PID 3936) and two
cmd.exe processes fan out to systeminfo.exe, net.exe → net1.exe,
whoami.exe and klist.exe (×2) — the full Cozy Bear reconnaissance sweep, linked into one
investigable incident.
powershell.exe (PID 3936) spawning
whoami.exe and klist.exe — Kerberos ticket enumeration flagged at 87% confidence
(ALR-175473), mapped to T1003.002.Both trees trace back to powershell.exe at PID 3936 — the same beacon-spawned shell the planner drove throughout
the operation. That shared root is the point: discovery and credential access are not two unrelated alerts, they are two
branches of one intrusion, and THRONE shows them that way.
Driving a live, autonomous adversary is a different test from replaying a script. The lesson is in what made it hard — and in what THRONE needed to keep up with a target that improvised.
APT29 is hard precisely because almost every individual action is plausible on a Windows host. whoami,
net, systeminfo, PowerShell, csc.exe, mstsc — administrators run all of
these every day. The malice is in the sequence and the causality: a shell that reads LSASS, then enumerates
tickets, then probes domain trusts, then reaches for RDP. A detection strategy built only on atomic signatures raises a
scatter of low-confidence alerts and misses the narrative that ties them together.
Three things carried the detection here. First, behavior over signature: INC-55853 caught LSASS access by what it did, so a renamed or custom credential tool would not have slipped past. Second, causality: rebuilding process trees from Sysmon GUIDs turned a wave of atomic alerts into a few investigable incidents with a common root. Third, autonomous triage: ONYX scored and grouped every alert the instant it landed, so the first thing a human would see is a ranked set of incidents — CRITICAL credential theft at the top — not an undifferentiated firehose.
Those four numbers describe the emulation library and THRONE's rule base — not a claim of coverage. We prove detections the only honest way: by driving real adversary behaviour at the platform and measuring what it catches, profile by profile, rather than asserting coverage on a slide. APT29 was one such profile, run end to end, measured as it happened.
Four concrete moves this operation argues for — whether or not you run THRONE.
Alert on any non-system process opening a read handle to lsass.exe (Sysmon ProcessAccess, or your EDR's
equivalent), not just on the string "mimikatz." APT29 has many ways to reach credential memory; the behavior is the constant.
Then raise the cost of success: enable LSA protection (RunAsPPL) and, where you can, Credential Guard, so the
memory is harder to reach in the first place.
csc.exe or MSBuild spawned by a script host, PowerShell reading from an alternate data stream, and Win32 API
calls from a command line are rarely benign in sequence. Turn on PowerShell script-block and module logging plus AMSI, and
score these higher when their parent is a shell or a freshly spawned beacon rather than an interactive admin session.
Individually benign discovery commands only become an intrusion when you can see they share a parent. Invest in process-tree / causality correlation — it is what collapsed eight-plus atomic alerts into a handful of incidents here, and it is what gives an analyst a story to follow instead of a queue to clear.
On an endpoint they steal credentials; in your tenant they abuse them. Enforce phishing-resistant MFA, audit OAuth application consents and service-principal credentials, watch for anomalous or forged tokens (the SolarWinds and Midnight Blizzard lessons), and gate RDP / SMB lateral movement behind privileged-access controls. The endpoint catch is where you stop them cheaply; the identity controls are where you stop them for good.
Every recommendation points the same way: stop grading single events and start reading the story they tell together. That is the gap a patient, living-off-the-land adversary like APT29 is built to exploit — and the gap THRONE is built to close.