Turn a business outcome into governed device operations
What happens if you take the automation model behind telecom's autonomous networks and apply it to enterprise device management? FleetOps is an add-on layer that sits beside an existing MDM/UEM platform. A warehouse doesn't care that a scanner app restarted; it cares that scanning kept working through the shift. So FleetOps takes an outcome like “scanner availability ≥ 99% during staffed shifts” and works toward it through the platform's supported interfaces only — with every action checked, bounded, verified and auditable.
The idea behind it
concepts borrowed, not a standards claimBeside the MDM/UEM platform a company already runs, not instead of it. The platform keeps doing what it does — enrolling devices, pushing profiles, sending commands. FleetOps adds the layer above: what outcome the business needs, whether it's being met, and which safe action to take next. It only acts through the platform's supported APIs.
Telecom operators run networks with an SMO (Service Management and Orchestration) layer and follow TM Forum's intent-based, autonomous-network model: state the outcome, measure it continuously, let pre-approved automation close the loop, and earn autonomy step by step. Device fleets have the same shape of problem — thousands of endpoints, strict uptime needs, changes that can hurt people's work — but rarely get that treatment.
That the pattern transfers. Everything here is synthetic, but each piece maps to an approach already used in network operations.
| Telecom networks | FleetOps | |
|---|---|---|
| Intent-based operations (TM Forum) | → | A business goal: scanners usable ≥ 99% of the shift |
| Service assurance | → | Availability measured in device-minutes, unknown counted |
| SMO / Non-RT RIC running rApps | → | The control layer running approved runbooks |
| Conflict management between rApps | → | Policy arbitration: security deadline vs. busy shift |
| Closed-loop automation | → | Detect → fix → verify, with human approval |
| Network digital twin | → | Fleet digital twin and rollout simulation |
| Southbound interfaces to network elements | → | Adapters to the UEM and diagnostics platforms |
| Autonomy levels L0–L5 (TM Forum) | → | Autonomy earned per task and device group |
Logical architecture — five planes, one execution gateway
Click a block for its job. Pulses show observations flowing up and approved actions flowing down.The numbering above is the build order, not a menu: all four share one intent model, one evidence store and one gateway. Four separate automation engines would mean four competing writers.
Design principles
- Outcomes are measurable. Every intent has an owner, population, metric, target, window, constraints and an assurance policy.
- Permission is explicit. Every action — AI-suggested or operator-initiated — passes the same gateway.
- Safety rules are deterministic. Models can propose and rank. They can't weaken a hard constraint or grant themselves authority.
- Uncertainty is visible. Stale or missing data yields unknown, never a healthy result.
- State is versioned. Intent, policy, snapshot, model and plan versions make any decision reproducible.
- Learning is governed. New models and runbooks are evaluated offline and approved before they touch production authority.
- Integration is replaceable. Product adapters isolate release differences, rate limits and auth quirks from planning.
“Keep warehouse scanners above 99% availability” isn't executable yet
Which scanners? Which shifts? What counts as usable? How are missing observations counted? Who can accept an exception? Compile the sentence and watch what's resolved from governed catalogs, what's inferred from templates, and what stays open for the owner.
Intent request
example tenantMachine-readable intent · revision 3
FleetOps format — not a vendor APILifecycle & assurance
Assurance is a separate axis: an active intent can be violated without being retired. Changes create a new revision and re-run admission; past evidence is never rewritten.
Availability semantics — what 99% actually allows
Availability = confirmed usable eligible minutes ÷ all eligible minutes. Unknown minutes stay in the denominator.Cap = lower of the absolute cap and 1% of the currently eligible population, rounded down. A result of zero means no automatic disruption at all. A fleet average can hide a full outage at a small site, so the intent also carries a site floor. These are starting hypotheses to validate, not general safety limits.
300 scanners across two distribution centres
Each tile is one enrolled scanner. The twin keeps what you asked for (desired), what the devices told you (reported) and what the system believes (inferred, with confidence and expiry) as three separate layers. A device that accepted an update command still shows the old version until fresh evidence arrives.
Shift-to-date measurement
—Snapshots
immutable manifestsCohorts — material attributes the simulator partitions on
synthetic OEMs & builds| Cohort | OEM | OS build | App | Devices | Qualified actions | Status |
|---|
Operational key = tenant + platform instance + enrollment ID + lifecycle epoch, so a re-enrolled device doesn't inherit someone else's history. Hardware serials help link history but aren't the sole authority.
Upgrade the picking app on 2,000 scanners
Aggregate forecasts are where rollouts go wrong: a healthy 95% can hide 100 devices on a build nobody has tested. Run the preflight and watch it return a conditional result instead of a green light.
Plan: pick-app 4.2.7 → 4.3.0 · 2,000 scanners · 6 sites
Cohort outcome
estimates, not guaranteesVerdict
Staged rollout with stop conditions
A small canary catches gross incompatibility. It can't prove a rare failure rate is acceptable — and correlated fleet outcomes make it weaker still.
Security patch due 18:00 vs. a warehouse that can't stop scanning
Security says patch the picking app by 18:00. Operations says never restart a device mid-transaction and never take more than 5 down at once. The scheduler packs patches into idle slots. Move the sliders until it can't — and see what it does instead of quietly lowering the security bar.
Patch schedule · 14:00 → 18:00 deadline
—Applicable policies · bundle v14
Decision
decision —Keeping controllers from fighting each other
One restart per device per 30 min, two remediation attempts per incident in the pilot. Prevents flapping between two configurations as metrics wobble.
Enter “at risk” at 99.0%, leave only above 99.2%. A single threshold turns noise into action.
Every managed profile field has one owning controller. A manual change in the UEM console invalidates affected plans; locks here can't stop another admin there, so readback matters.
12 scanners at DC-01 stop passing the app health check
300 scanners are covered by the readiness intent. Inventory and connectivity are fresh, but 12 devices fail an authorised synthetic transaction. Step through detection, diagnosis, planning, approval, bounded execution and independent verification.
Durable workflow
idleAffected devices
cap = min(5, ⌊1% × 300⌋) = 3 per batchDiagnosis — competing hypotheses
ranked, not provenWaiting for a persistent deviation.
Approval
—No plan awaiting approval.
What's happening · live log
newest at the bottom✓ done / passed · ✕ blocked · ! needs attention · ▶ change sent to devices · • information. Small grey text is the technical detail for engineers.
One guarded door between a decision and a device
Every command passes through the same checkpoint. It re-checks permission at the last second, writes down what it's about to do, and — if a reply goes missing — finds out what actually happened before trying again. The goal: a scanner never gets restarted twice by accident.
How one command reaches a scanner — step by step
Idempotency key
If the platform supports an idempotency key, pass it. If not, rely on local dedup plus task lookup or device readback. A timeout after submission is ambiguous: the device may already have acted.
Event envelope
eventId evt_01J… eventType execution.stepChanged schemaVersion 1 tenantId customer-a productInstanceId uem-east-1 deviceInstanceId SCN-0147@3 observedAt 2026-09-23T11:42:07Z receivedAt 2026-09-23T11:42:09Z correlationId inc-2291 causationId dec-5530 quality ["fresh"]
Control-layer API
proposed · custom, not TMF921-conformant| Method & path | Contract behaviour |
|---|---|
| POST /v1/intents | Draft; validates tenant & schema |
| POST /v1/intents/{id}/admissions | Async report; no device changes |
| GET /v1/intents/{id}/assurance | Window, coverage, freshness, reasons |
| POST /v1/simulations | 202 + resource location |
| POST /v1/decisions | Decision ref + disposition |
| POST /v1/approvals | Approver role, expiry, plan hash |
| POST /v1/executions | Idempotency key + revalidation |
| GET /v1/executions/{id} | Accepted ≠ applied ≠ verified |
| POST /v1/executions/{id}/stop | Stopping state + in-flight limits |
| GET /v1/evidence/{id} | Tenant, role, classification filters |
409 for conflicting revisions or execution state, 422 for semantic validation failures, 429 with retry guidance. ETag / If-Match on intent updates.
Autonomy is earned one scenario at a time
The same system can run app recovery at level 3 while OS changes stay at level 1. These levels are an adaptation inspired by the TM Forum autonomous-networks progression — not a certification.
Watch the same morning at level 2, 3 and 4
At a glance
| Level 2 · approve each plan | Level 3 · pre-approved limits | Level 4 · coordinated | |
|---|---|---|---|
| Who says yes to a fix? | A person, every single time | The owner, once, by signing off the limits in advance | Same as level 3, plus agreements with other teams |
| What people still do | Review and approve every plan; handle leftovers | Handle exceptions; read the summary | Set goals and limits; do physical tasks; supervise results |
| What happens outside the limits | — | It stops and asks a person | It stops and asks; other teams' systems decide for themselves |
| Status here | first release (MVP) | next, after evidence | illustrative, later |
Operating levels
Scenario matrix
Promotion evidence · app restart → L3
A lack of incidents isn't evidence if few actions were attempted. Demotion is immediate on broken isolation, an unauthorised action or loss of required verification.
The happy path proves nothing. Run the drills.
Each drill injects a condition the MVP has to survive. Pass means the system did the safe thing — which is often nothing, loudly.
Drills
Expected safe behaviour
Pick a drill.
✓ done / passed · ✕ blocked · ! needs attention · ▶ change sent to devices · • information. Small grey text is the technical detail for engineers.
One customer, two sites, 200–500 Android scanners, two reversible actions
Core team of 8–10 with shared security, platform, OEM and customer operations support. Don't fund a full-fidelity device emulator first: a calibrated fleet-state model, dependency validation and well-instrumented canaries pay back sooner.
Phases
planning sequence, not a commitmentAcceptance gates
| Gate | Proposed criterion |
|---|---|
| Integration | All required reads and actions pass on the qualified release & OEM matrix |
| Measurement | ≥ 98% eligible-minute coverage for 10 consecutive operating days |
| Decision safety | Every dispatched action has a valid policy decision and exact-scope authority |
| Simulation | All seeded incompatible cases rejected; unknown cohorts flagged |
| Recovery | Stop, ambiguous-timeout and restart tests pass with no unsafe duplicates |
| Outcome | ≥ 90% of qualified recoveries independently verified (min. 50 attempts) |
| Harm | No unauthorised or severe harmful action |
| Operations | Customer support can run and stop the loop from documented runbooks |
KPIs that keep it honest
| KPI | Why it's there |
|---|---|
| Verified recovery time | Median and p95 vs. baseline. Faster dispatch with unchanged downtime isn't success. |
| Harmful action rate | Severity-classified; severe harm triggers stop. |
| False-safe simulation rate | Plans rated safe that caused a predicted harm category. |
| Unnecessary action rate | Blinded review where practical. |
| Operator effort | Approval and exception work counts, not just saved work. |
| Audit completeness | 100% required for pilot dispatch. |
| Cost per verified recovery | Includes inference, storage, support and failed attempts. |
Illustrative business target: 20% lower median verified recovery time for qualified incident types, with no worse p95 or harm rate. Final target set after baseline collection.
Top risks
| Risk | Impact | Mitigation · owner |
|---|---|---|
| Required API fields or actions unavailable | Scenario can't execute | Capability spike first; supported fallback or narrower scope · integration lead |
| Business availability poorly defined | Optimises device proxies while operations still fail | Signed metric contract + business telemetry · service owner |
| Twin overstates certainty | Unsafe rollout accepted | Explicit coverage, abstention, lab tests, canaries · simulation lead |
| Delayed commands execute in an unsafe window | Interruption after context changed | Execution-time guards or exclude from automatic use · integration lead |
| Approval fatigue | Rubber-stamping | Risk tiering, compact evidence, measured approval workload · product owner |
| Model or retrieval compromise | Unauthorised recommendations | No tool credentials in the model path; schema validation; deterministic gates · security owner |
What's real here and what isn't
Scope
Kaldrivon FleetOps is a browser-only reference demo of an architecture for outcome-driven device operations: intent management, a fleet digital twin, policy arbitration and closed-loop remediation sharing one governed execution path. It's written from a practitioner's view of how to build it safely, not a product pitch.
It's deliberately vendor-neutral. “UEM platform” and “Diagnostics platform” stand in for whatever device management and diagnostics products a customer already runs. No vendor endpoint paths are invented; every adapter operation would be qualified against the installed release during discovery.
What's synthetic
Everything. Devices, OEM names, OS builds, telemetry, hypotheses, probabilities, schedules and events are generated in your browser with a fixed seed. Thresholds (99%, 98% coverage, 5-device cap, 30-minute cooldown, 120-second freshness) are starting hypotheses from the design, not validated safety limits. Nothing leaves this page.
Public references
- TM Forum — Intent-based automation · intent owner / handler framing
- TMF921 Intent Management API repository · interoperability direction only; no conformance claimed
- TM Forum — Autonomous Networks · level taxonomy the mobility levels adapt
Related Kaldrivon demos
- Kaldrivon SMO · O-RAN SMO / Non-RT RIC reference implementation
- Kaldrivon O1 Simulator · NETCONF/YANG fault & PM
- Kaldrivon PulseHub