Bareket AIPowerGuardBAREKET AIRequest a pilot
Home|Technology
TECHNOLOGY

What we read, and what we defend against.

A compromised server can falsify every log it writes. It cannot falsify its own power draw. This page follows that signal end to end: the buses we read, the detectors that judge them, the hardware it is proven on, and exactly what leaves your environment. Written for the people who run the infrastructure.

Three tiers. Yours, at every one of them.

The analysis tier is installed on your infrastructure, not inside your servers. And phase 01 installs nothing anywhere.

01

On the server

From phase 02, the PowerGuard agent runs at the management layer of each server, sampling the power rails continuously and seeing the command layer directly.

02

On the edge

An appliance on your network aggregates fleet telemetry and runs the analysis. In phase 01 it is the only component involved, reading over the standard management API, installing nothing.

03

Your cloud, or fully local

The fleet view runs in your cloud tenancy or entirely on premises. Air-gapped operation is a supported configuration, not a special case.

Two collection modes, one picture.

PowerGuard reads the power-control plane of the server: the BMC, the PMBus bus, the voltage regulators and the rails they drive. Collection runs in one of two modes, and the honest comparison between them matters.

CapabilityManagement APInothing installed, phase 01On-server agentphases 02–03
Inventory & topologyFull fleetFull fleet
Firmware & configuration stateRead on every pollRead continuously
Telemetry samplingPeriodic, seconds per pollContinuous: 100 Hz per controller
Command-layer visibilityNot visibleEvery power command observed
Baseline resolutionCoarse, drift and configurationFine, millisecond physical behaviour
ResponseReport onlyBlock under your signed policy, phase 03

Layered detectors, one verdict.

No single test decides anything. Independent detectors read the same telemetry, and enforcement acts only on physical or protocol evidence, never on a statistical anomaly alone.

Envelope

Every reading is checked against the physical operating envelope learned for that exact unit, not a datasheet, your hardware.

Change-point

Abrupt shifts in behaviour are flagged the moment the statistics of the signal change, before thresholds are ever crossed.

Desynchronisation

Commands and physical response are compared continuously. Effect without cause, or cause without effect, is evidence.

Statistical distance

Each server's behaviour is measured against its own history and the fleet's, so slow drift stands out as clearly as a spike.

Consistency witness

Independent readings of the same physical quantity are cross-checked, so a lying sensor cannot pass unnoticed.

Behavioural, not generative.

There is no language model in the detection path. PowerGuard learns the physical behaviour of your own hardware and measures how far the present sits from it.

Learned from your fleet

The baseline comes from your hardware under your workload, not from a generic profile.

It learns only when things are normal

The model does not absorb an attack in progress into its own definition of normal.

Distance, not opinion

Deviation is measured, and every verdict carries the measurement behind it.

Nothing leaves

The model runs inside your environment; your telemetry does not train anything for anyone else.

“Your alerts will fire every time we start a training run.” No, and here is why.

A coordinated AI workload swing looks dramatic on a power chart. Structurally, it is nothing like protocol abuse, and the difference is measurable on four independent axes.

Shape

Legitimate load moves inside the physical envelope learned for that exact unit, ramps the hardware was designed for. Destructive commands push rails outside that envelope, which is precisely what the envelope detector exists to see.

Timing

Workload transitions follow scheduler ramps with characteristic rise profiles. Attacks produce discontinuities the physics of a healthy system does not produce on its own.

Correlation

A training run moves hundreds of servers in step, the fleet view shows one coherent event. An attack rarely does; a single server departing from a fleet-wide pattern is signal, not noise.

Cause

Every legitimate physical change is preceded by a command chain that explains it. Physical effect with no explaining command, or a command with no matching effect, is evidence, and that is the desynchronisation detector's whole job.

A legitimate load spike is examined, never punished: response requires physical or protocol evidence, and the record-only phase exists so you can watch every decision the system would have made on your own workload, including your training runs, before enforcement is ever enabled.

One engine. It sees a failure coming, and an attack landing.

The baseline that judges every reading against the physics of a healthy unit does not care why a unit departs from it. Slow drift means ageing hardware. A millisecond discontinuity means something else entirely. Both come out of the same measurement.

EVERY DAY · FOR OPERATIONS

Every controller carries a health score derived from its own baseline. When a component starts to drift, you get 24–72 hours of warning, and a ticket in your queue before the outage, not a post-mortem after it.

  • Health score per power controller, recalculated continuously
  • Failure predicted 24–72 hours ahead
  • Ticket raised automatically in the tools you already run

Early-warning capability is being validated with design partners on production hardware.

THE DAY IT MATTERS · FOR SECURITY

The same engine reads a departure measured in milliseconds as what it is: a command doing something the physics of a healthy system never does on its own, detected at 20 ms in lab validation, before silicon damage.

  • Attack signatures and degradation read from one baseline
  • PMFault-class events detected before silicon damage
  • No second agent, no second deployment

A voltage attack.
Detected before
silicon damage.

A live attack drove a rail from 5.21 V to 4.64 V. PowerGuard identified the physical evidence and produced a verdict in 20 milliseconds.

REAL HARDWARELIVE ATTACKREPEATABLE TEST

The attack class is public research: PMFault (TCHES 2023) showed that PMBus overvoltage commands can permanently destroy server CPUs with no physical access. Our Gen 0 testbed reproduces the class of event, and detects it from the physics, at 100 Hz, before damage.

Turn invisible infrastructure into actionable context.

Understand every rail, controller and relationship. Then move from anomaly to decision with evidence your security and operations teams can trust.

  • Continuous asset discovery
  • Physics-based attack detection
  • Three-state response logic
  • Compliance-ready evidence
POWERGUARD / ASSET GRAPH LIVE
POWER STATENOMINALACTIVE RAILS24 / 24SCAN RATE100 Hz
BMC-01
PMBUS
POWERGUARD
VRM-04
GPU RACK
No destructive commands detected Updated 2s ago

Not a mockup.
Streaming right now.

PowerGuard metrics from real hardware, live in Datadog: rail voltage and current, anomaly score, alert status and thermal state. This is the actual security dashboard our Gen 0 unit reports to.

app.datadoghq.eu / powerguard-security-dashboard LIVE
PowerGuard Security Dashboard in Datadog: system overview tiles with alert status, anomaly score, output current, output voltage and temperature, plus anomaly-score and alert-status charts over time
Datadog widgets: anomaly score over time with critical threshold, and alert status over time
SECURITY & ANOMALY DETECTION
Datadog widgets: output current and output voltage over time
ELECTRICAL MONITORING · IOUT / VOUT
Datadog widgets: temperature over time with warning and critical thresholds table
THERMAL MONITORING & THRESHOLDS

Two surfaces. One source of truth.

The console is the daily-use interface. The same findings publish into the security and operations stack you already run, you choose which surface each team lives in.

SIEM

Findings and evidence delivered into Splunk and Microsoft Sentinel, in the format your correlation rules already expect.

Observability

Health and security metrics stream to Datadog, the live dashboard above is exactly this integration, running.

Operations

Predictive findings raise tickets in ServiceNow automatically, so degrading hardware enters the queue your team already works.

Splunk SOC: PowerGuard confirmed verdicts delivered into the SIEM, multi-layer and high severity
Splunk SOC · the SIEM integration, running on our bench

Measured at the edge. Forwarded as findings.

Continuous power telemetry at fleet scale is a volume no northbound platform should ingest raw. The pipeline is built so it never has to.

01

Sampled

Up to 100 Hz per power controller in agent mode; periodic polls over the management API in phase 01. Raw samples never leave the tier that produced them.

02

Analysed

The layered detectors run at the edge, on your infrastructure. Every sample is examined; almost none needs to travel.

03

Retained

Full-resolution history stays local, at a retention you set, available for forensics and baselining for as long as you want it, and no longer.

04

Forwarded

Verdicts, findings and evidence packs go northbound to the tools you already run, orders of magnitude smaller than the raw stream they summarise.

The exact reduction ratio depends on fleet size and sampling mode; it is measured on your own fleet during the pilot and written into the success criteria, a factual number, not a brochure claim.

Evidence your auditor can hold.

The power layer is entering the regulations · PMBus 1.5 on the protocol side, NIS2 and IEC 62443 on the operator side. PowerGuard assesses your fleet against them continuously, not the week before the audit.

PMBus 1.5

Configuration and command posture assessed against the protocol's own secure-device profile, unit by unit.

NIS2 & IEC 62443

The same measurements mapped to the controls your regulators and customers audit against.

Evidence packs

Findings exported as documents an auditor accepts, generated from your own fleet, on demand.

Built for the environments that stay isolated.

Run locally. Keep control. The home page states the position, this is where a security architect verifies it: what runs where, what leaves the environment, and who holds authority.

01Local processing, detection and response run locally, air-gapped and disconnected operation supported02Machine telemetry only: voltage, current, temperature, power and management-plane commands; no workload content03Nothing has to leave, the analysis tier deploys on-premises or in your cloud tenancy; northbound forwarding is optional and under your control04Policy and keys live in your environment, enforcement acts only under policy you sign, and that authority is yours to revoke at any time

The standard, extended up the supply chain.

Your fleet is measured against a physical standard. PowerCert offers that same standard to the vendors who make the power components themselves, so the parts in your next servers arrive already validated against the detectors that will watch them in production.

One standard

Components are exercised against the same envelope, timing and command detectors that run in the field, certification and production measure the same physics.

Silicon & GaN

The program covers the components that actually switch the power, silicon MOSFET and GaN alike, certified by the vendors who build them.

Certified before it ships

A certified part means its behaviour under attack conditions is known before it enters your fleet, not discovered there.

The public record this work stands on.

Stated plainly.

  • Validation to date is on our own instrumented hardware. Behaviour on production power hardware is being validated with design partners.
  • We detect the presence of cryptographic activity in laboratory conditions. We do not recover keys, and we do not claim to.
  • We do not defend against an attacker who already holds root on the server management controller.
  • Early-warning capability for hardware failure is in validation and is not yet a measured product guarantee.