KaldrivonFleetOps
Intent-driven device operations
● Synthetic data DC-01 · shift 06:00–14:00 L2 · execute on approval
Enterprise mobility · autonomous device operations

Turn a business outcome into governed device operations

What happens if you take the automation model behind telecom's autonomous networks and apply it to enterprise device management? FleetOps is an add-on layer that sits beside an existing MDM/UEM platform. A warehouse doesn't care that a scanner app restarted; it cares that scanning kept working through the shift. So FleetOps takes an outcome like “scanner availability ≥ 99% during staffed shifts” and works toward it through the platform's supported interfaces only — with every action checked, bounded, verified and auditable.

The idea behind it

concepts borrowed, not a standards claim
Where it sits

Beside the MDM/UEM platform a company already runs, not instead of it. The platform keeps doing what it does — enrolling devices, pushing profiles, sending commands. FleetOps adds the layer above: what outcome the business needs, whether it's being met, and which safe action to take next. It only acts through the platform's supported APIs.

Where the idea comes from

Telecom operators run networks with an SMO (Service Management and Orchestration) layer and follow TM Forum's intent-based, autonomous-network model: state the outcome, measure it continuously, let pre-approved automation close the loop, and earn autonomy step by step. Device fleets have the same shape of problem — thousands of endpoints, strict uptime needs, changes that can hurt people's work — but rarely get that treatment.

What this demo shows

That the pattern transfers. Everything here is synthetic, but each piece maps to an approach already used in network operations.

Telecom networksFleetOps
Intent-based operations (TM Forum)→A business goal: scanners usable ≥ 99% of the shift
Service assurance→Availability measured in device-minutes, unknown counted
SMO / Non-RT RIC running rApps→The control layer running approved runbooks
Conflict management between rApps→Policy arbitration: security deadline vs. busy shift
Closed-loop automation→Detect → fix → verify, with human approval
Network digital twin→Fleet digital twin and rollout simulation
Southbound interfaces to network elements→Adapters to the UEM and diagnostics platforms
Autonomy levels L0–L5 (TM Forum)→Autonomy earned per task and device group
Service availability—Intent target ≥ 99.00% · shift to date
Observation coverage—Pilot floor 98% · unknown counted
Error budget left—non-usable device-minutes
Open incidents0grouped by site · service · time
Assurance state—satisfied · at risk · violated · unknown

Logical architecture — five planes, one execution gateway

Click a block for its job. Pulses show observations flowing up and approved actions flowing down.
Governance wraps every plane: tenant isolation, scoped credentials, approval expiry, audit, action budgets, stop controls and model release review. The execution gateway is the only new component allowed to dispatch production changes.

The numbering above is the build order, not a menu: all four share one intent model, one evidence store and one gateway. Four separate automation engines would mean four competing writers.

Design principles

  • Outcomes are measurable. Every intent has an owner, population, metric, target, window, constraints and an assurance policy.
  • Permission is explicit. Every action — AI-suggested or operator-initiated — passes the same gateway.
  • Safety rules are deterministic. Models can propose and rank. They can't weaken a hard constraint or grant themselves authority.
  • Uncertainty is visible. Stale or missing data yields unknown, never a healthy result.
  • State is versioned. Intent, policy, snapshot, model and plan versions make any decision reproducible.
  • Learning is governed. New models and runbooks are evaluated offline and approved before they touch production authority.
  • Integration is replaceable. Product adapters isolate release differences, rate limits and auth quirks from planning.
Intent Studio · from a sentence to an enforceable objective

“Keep warehouse scanners above 99% availability” isn't executable yet

Which scanners? Which shifts? What counts as usable? How are missing observations counted? Who can accept an exception? Compile the sentence and watch what's resolved from governed catalogs, what's inferred from templates, and what stays open for the owner.

Intent request

example tenant
Or try:
NOT EVALUATEDCompile to run admission. No device changes happen during admission.

Machine-readable intent · revision 3

FleetOps format — not a vendor API


      

Lifecycle & assurance

Assurance is a separate axis: an active intent can be violated without being retired. Changes create a new revision and re-run admission; past evidence is never rewritten.

Availability semantics — what 99% actually allows

Availability = confirmed usable eligible minutes ÷ all eligible minutes. Unknown minutes stay in the denominator.
Eligible device-minutes
480,000
Non-usable allowance
4,800
Per-device average
4.8 min
Effective concurrency cap
5

Cap = lower of the absolute cap and 1% of the currently eligible population, rounded down. A result of zero means no automatic disruption at all. A fleet average can hide a full outage at a small site, so the intent also carries a site floor. These are starting hypotheses to validate, not general safety limits.

Fleet digital twin · versioned state, not a device emulator

300 scanners across two distribution centres

Each tile is one enrolled scanner. The twin keeps what you asked for (desired), what the devices told you (reported) and what the system believes (inferred, with confidence and expiry) as three separate layers. A device that accepted an update command still shows the old version until fresh evidence arrives.

Shift-to-date measurement

—
Confirmed usable—
Known unavailable—
Unknown (in denominator)—
Availability (conservative)—
Observation coverage—
Worst site—

Snapshots

immutable manifests
No snapshot yet. Every simulation and decision pins one.

Cohorts — material attributes the simulator partitions on

synthetic OEMs & builds
CohortOEMOS buildAppDevicesQualified actionsStatus

Operational key = tenant + platform instance + enrollment ID + lifecycle epoch, so a re-enrolled device doesn't inherit someone else's history. Hardware serials help link history but aren't the sole authority.

Simulation & rollout prediction · progressive fidelity

Upgrade the picking app on 2,000 scanners

Aggregate forecasts are where rollouts go wrong: a healthy 95% can hide 100 devices on a build nobody has tested. Run the preflight and watch it return a conditional result instead of a green light.

Plan: pick-app 4.2.7 → 4.3.0 · 2,000 scanners · 6 sites

Cohort outcome

estimates, not guarantees

Verdict

—
Run the simulation. Possible results: pass · conditional · fail · insufficient evidence.

Staged rollout with stop conditions

Each wave waits for its observation window, service-health evidence and approval policy. The rollout stops automatically if unknown observations exceed the allowed coverage gap — even when the observations that did arrive look healthy.
Canary size with zero failures20
95% upper bound on failure rate13.9%
Rule-of-three approximation (3/n)15.0%

A small canary catches gross incompatibility. It can't prove a rare failure rate is acceptable — and correlated fleet outcomes make it weaker still.

Policy arbitration · mandatory constraints before preferences

Security patch due 18:00 vs. a warehouse that can't stop scanning

Security says patch the picking app by 18:00. Operations says never restart a device mid-transaction and never take more than 5 down at once. The scheduler packs patches into idle slots. Move the sliders until it can't — and see what it does instead of quietly lowering the security bar.

Patch schedule · 14:00 → 18:00 deadline

—
Patches scheduled in slot Idle capacity (no active transaction) Spare-swap capacity Missed deadline

Applicable policies · bundle v14

Decision

decision —

Keeping controllers from fighting each other

Dwell time & cooldown

One restart per device per 30 min, two remediation attempts per incident in the pilot. Prevents flapping between two configurations as metrics wobble.

Hysteresis

Enter “at risk” at 99.0%, leave only above 99.2%. A single threshold turns noise into action.

Field ownership

Every managed profile field has one owning controller. A manual change in the UEM console invalidates affected plans; locks here can't stop another admin there, so readback matters.

Closed-loop remediation · scenario: app stalls mid-shift

12 scanners at DC-01 stop passing the app health check

300 scanners are covered by the readiness intent. Inventory and connectivity are fresh, but 12 devices fail an authorised synthetic transaction. Step through detection, diagnosis, planning, approval, bounded execution and independent verification.

Durable workflow

idle

Affected devices

cap = min(5, ⌊1% × 300⌋) = 3 per batch

Diagnosis — competing hypotheses

ranked, not proven

Waiting for a persistent deviation.

Approval

—

No plan awaiting approval.

What's happening · live log

newest at the bottom

✓ done / passed · ✕ blocked · ! needs attention · ▶ change sent to devices · • information. Small grey text is the technical detail for engineers.

Safe execution · the only path to a device

One guarded door between a decision and a device

Every command passes through the same checkpoint. It re-checks permission at the last second, writes down what it's about to do, and — if a reply goes missing — finds out what actually happened before trying again. The goal: a scanner never gets restarted twice by accident.

How one command reaches a scanner — step by step

a real change on its waya check or a questionOK / confirmedrefused or lostcaution
    Pick a case to replay it.

    Idempotency key

    exec-7f3a+dev epoch 3+step 2+desired v18→idk_9c41e2…

    If the platform supports an idempotency key, pass it. If not, rely on local dedup plus task lookup or device readback. A timeout after submission is ambiguous: the device may already have acted.

    Event envelope

    eventId        evt_01J…
    eventType      execution.stepChanged
    schemaVersion  1
    tenantId       customer-a
    productInstanceId uem-east-1
    deviceInstanceId  SCN-0147@3
    observedAt     2026-09-23T11:42:07Z
    receivedAt     2026-09-23T11:42:09Z
    correlationId  inc-2291
    causationId    dec-5530
    quality        ["fresh"]

    Control-layer API

    proposed · custom, not TMF921-conformant
    Method & pathContract behaviour
    POST /v1/intentsDraft; validates tenant & schema
    POST /v1/intents/{id}/admissionsAsync report; no device changes
    GET /v1/intents/{id}/assuranceWindow, coverage, freshness, reasons
    POST /v1/simulations202 + resource location
    POST /v1/decisionsDecision ref + disposition
    POST /v1/approvalsApprover role, expiry, plan hash
    POST /v1/executionsIdempotency key + revalidation
    GET /v1/executions/{id}Accepted ≠ applied ≠ verified
    POST /v1/executions/{id}/stopStopping state + in-flight limits
    GET /v1/evidence/{id}Tenant, role, classification filters

    409 for conflicting revisions or execution state, 422 for semantic validation failures, 429 with retry guidance. ETag / If-Match on intent updates.

    Human controls · autonomy per scenario and cohort

    Autonomy is earned one scenario at a time

    The same system can run app recovery at level 3 while OS changes stay at level 1. These levels are an adaptation inspired by the TM Forum autonomous-networks progression — not a certification.

    Watch the same morning at level 2, 3 and 4

    Press Play to watch it step by step.

    At a glance

    Level 2 · approve each planLevel 3 · pre-approved limitsLevel 4 · coordinated
    Who says yes to a fix?A person, every single timeThe owner, once, by signing off the limits in advanceSame as level 3, plus agreements with other teams
    What people still doReview and approve every plan; handle leftoversHandle exceptions; read the summarySet goals and limits; do physical tasks; supervise results
    What happens outside the limits—It stops and asks a personIt stops and asks; other teams' systems decide for themselves
    Status herefirst release (MVP)next, after evidenceillustrative, later

    Operating levels

    Scenario matrix

    •current•next candidate•demoted

    Promotion evidence · app restart → L3

      A lack of incidents isn't evidence if few actions were attempted. Demotion is immediate on broken isolation, an unauthorised action or loss of required verification.

      Demonstration sequence · prove controlled behaviour under failure

      The happy path proves nothing. Run the drills.

      Each drill injects a condition the MVP has to survive. Pass means the system did the safe thing — which is often nothing, loudly.

      Drills

      Expected safe behaviour

      Pick a drill.

      ✓ done / passed · ✕ blocked · ! needs attention · ▶ change sent to devices · • information. Small grey text is the technical detail for engineers.

      MVP · indicative 16 weeks after the integration feasibility gate

      One customer, two sites, 200–500 Android scanners, two reversible actions

      Core team of 8–10 with shared security, platform, OEM and customer operations support. Don't fund a full-fidelity device emulator first: a calibrated fleet-state model, dependency validation and well-instrumented canaries pay back sooner.

      Phases

      planning sequence, not a commitment

      Acceptance gates

      GateProposed criterion
      IntegrationAll required reads and actions pass on the qualified release & OEM matrix
      Measurement≥ 98% eligible-minute coverage for 10 consecutive operating days
      Decision safetyEvery dispatched action has a valid policy decision and exact-scope authority
      SimulationAll seeded incompatible cases rejected; unknown cohorts flagged
      RecoveryStop, ambiguous-timeout and restart tests pass with no unsafe duplicates
      Outcome≥ 90% of qualified recoveries independently verified (min. 50 attempts)
      HarmNo unauthorised or severe harmful action
      OperationsCustomer support can run and stop the loop from documented runbooks

      KPIs that keep it honest

      KPIWhy it's there
      Verified recovery timeMedian and p95 vs. baseline. Faster dispatch with unchanged downtime isn't success.
      Harmful action rateSeverity-classified; severe harm triggers stop.
      False-safe simulation ratePlans rated safe that caused a predicted harm category.
      Unnecessary action rateBlinded review where practical.
      Operator effortApproval and exception work counts, not just saved work.
      Audit completeness100% required for pilot dispatch.
      Cost per verified recoveryIncludes inference, storage, support and failed attempts.

      Illustrative business target: 20% lower median verified recovery time for qualified incident types, with no worse p95 or harm rate. Final target set after baseline collection.

      Top risks

      RiskImpactMitigation · owner
      Required API fields or actions unavailableScenario can't executeCapability spike first; supported fallback or narrower scope · integration lead
      Business availability poorly definedOptimises device proxies while operations still failSigned metric contract + business telemetry · service owner
      Twin overstates certaintyUnsafe rollout acceptedExplicit coverage, abstention, lab tests, canaries · simulation lead
      Delayed commands execute in an unsafe windowInterruption after context changedExecution-time guards or exclude from automatic use · integration lead
      Approval fatigueRubber-stampingRisk tiering, compact evidence, measured approval workload · product owner
      Model or retrieval compromiseUnauthorised recommendationsNo tool credentials in the model path; schema validation; deterministic gates · security owner
      About this demo

      What's real here and what isn't

      Idea, architecture & demo

      Janos Korognai

      Founder & CTO of Kaldrivon. A telecom pre-sales and solution architect from Nokia's OSS and network management world — service assurance, closed-loop automation, intent-based operations and SMO / Non-RT RIC for Tier-1 North American operators. FleetOps takes the same ideas that keep radio networks running — measurable intent, assurance, governed closed loops — and applies them to enterprise device fleets.

      Scope

      Kaldrivon FleetOps is a browser-only reference demo of an architecture for outcome-driven device operations: intent management, a fleet digital twin, policy arbitration and closed-loop remediation sharing one governed execution path. It's written from a practitioner's view of how to build it safely, not a product pitch.

      It's deliberately vendor-neutral. “UEM platform” and “Diagnostics platform” stand in for whatever device management and diagnostics products a customer already runs. No vendor endpoint paths are invented; every adapter operation would be qualified against the installed release during discovery.

      What's synthetic

      Everything. Devices, OEM names, OS builds, telemetry, hypotheses, probabilities, schedules and events are generated in your browser with a fixed seed. Thresholds (99%, 98% coverage, 5-device cap, 30-minute cooldown, 120-second freshness) are starting hypotheses from the design, not validated safety limits. Nothing leaves this page.

      Public references

      Related Kaldrivon demos