Skip to main content

Documentation

No results found.
Features

Agency Fleet Dashboard

A read-only overview of every client install phoning home to this license server, at Dashboard → Fleet (dashboard/fleet; the old dashboard/memberships/fleet URL redirects — the fleet surface moved out of the Memberships product into its own...

A read-only overview of every client install phoning home to this license server, at Dashboard → Fleet (dashboard/fleet; the old dashboard/memberships/fleet URL redirects — the fleet surface moved out of the Memberships product into its own module, app/Features/Fleet).

Mothership/owner-facing. This page only exists on the license server itself (CMS_IS_LICENSE_SERVER=true — the webprocms.com mothership). On every other install the route 404s and the sidebar item is hidden. It is not a client-facing feature and is deliberately excluded from the member-facing docs importer.

What it shows

One row per install — a member holding a license_key. Everything on the page is already persisted by the license check-in flow (MembershipCheckController stamps last_seen_at, last_seen_host, and last_seen_version on every recognized check-in); the fleet page adds no schema and performs no writes.

Column Source
Host last_seen_host, with the member's email underneath; links to the member edit page
Version last_seen_version, plus an amber Update available → vX badge when it trails the latest published release (version_compare vs the release.version Setting written by cms:publish-release)
Last check-in last_seen_at as relative time — amber after 7 days, red after 30 days (stale), "Never" if the install has not checked in
Plan The active license-unlocking line item's plan (or Free), an Active/Inactive verdict (paid unlock OR comp/exempt, same signal the license check uses), and a Comp badge for exempt members
Notes admin_notes, truncated with the full text in a hover tooltip

Stats header

Five tiles: Installs (members with a license key), On latest, Update available (both computed against release.version; both read 0 until a release has been published), Stale (30d+) (checked in, but more than 30 days ago), and Comped (is_exempt).

Search & sorting

The search box matches host, license key, admin notes, or member email. Rows are sorted by most recent check-in first (never-seen installs last), paginated at 25.

Access control

Two gates, both of which must pass (updated 2026-07-24 — the feature:memberships gate was removed; see "v2 batch F" below):

  1. Membership::isFleetServer() — CMS_IS_LICENSE_SERVER=true (the mothership), or the Enterprise-tier Fleet Management feature enabled. Enforced by the EnsureFleetServer route middleware plus the components' own mount() guards; the sidebar item and the dashboard Fleet card render on the same verdict.
  2. Admin role or above (route middleware role:admin).

The consumer Memberships feature is deliberately not a gate — an agency runs a fleet without selling memberships.

Scope (v1)

Read-only by design: no write actions, no per-install update pushes, no deauthorization from this page. Manage the underlying member (comp flag, notes, license key) from the member edit page the host column links to. Publishing releases stays on cms:publish-release / the release runbook (docs/releasing.md).

v2 — Tabs, filters, drill-down, rollout & top errors (2026-07)

The page is now three tabs (URL-persisted via ?tab=):

  • Installs (default) — the v1 table, plus:
    • Clickable stat tiles as filters. Every tile (Installs / On latest / Update available / Stale / Comped / Erroring) filters the table; clicking the active tile clears the filter. Combines with search. Persisted via ?filter=.
    • Per-install drill-down. A chevron on the host cell expands the member row into its distinct installs (member_installs, one row per install UUID) — host, short install id, version with an Outdated badge, update speed, last check-in. This is where a member running www + staging + dev stops collapsing into one row.
    • Speed select. Each drill-down row has a Speed select: "Their choice" follows the channel the install's admin picked on their own Updates page, while picking Early / Standard / Cautious sets a fleet override (update_channel_override) that wins at serve time. Forcing Early makes the install a first-updater — offered each release immediately (this replaced the old canary flag, which was exactly fleet-forced Early); Cautious slows a fragile site down. (2026-07: the is_canary flag and cms:canary were removed; existing canaries were migrated to Early overrides.)
  • Rollout — release-health operations:
    • Red circuit-breaker banner when release.paused is set (auto-tripped by failing §D.4 health reports), with a confirm-modal Resume rollout button that clears the pause. Publishing a new release also resets it.
    • Current release card (version + published-at), Stability windows card (the channel soak hours), First updaters count (installs whose effective speed is Early). (The publish-time canary window was removed 2026-07 — staged rollout is the channel soak.)
    • Version adoption — per-version install counts with proportional bars; the latest version is badged.
    • Update reports — the most recent install_update_reports (post-apply health reports), failures first, each with status/probe badges, host, version, message, and age.
  • Top errors — install_error_reports grouped by signature across the whole fleet: class, file:line, sample message, distinct-installs-affected count (red when >1), total occurrences, the CMS versions it was seen on, and last-seen age. Sorted by installs affected — the "did this release break something everywhere" view.

Rollout and Top-errors data only load when their tab is open; the default Installs view stays as light as v1. All v2 actions (speed override, resume rollout) still sit behind the same three gates (license server + memberships feature + admin).

v2 batch B — health block, unlicensed installs, last IP (2026-07)

  • Health block. Every license check-in now carries a best-effort health block collected by app/Support/InstallHealth.php (client side): PHP version, scheduler-heartbeat age, free-disk %, and days until the site's TLS certificate expires (cached daily; null on http/failure/test env). Ingest whitelists + clamps the fields onto member_installs.health (plus php_version). The Fleet drill-down shows per-install badges — Cert expiring (≤14d), Cron dead (heartbeat >13h ≈ two missed polls), Disk low (<10% free), or OK — and the member row gets a red Health rollup badge when any of its installs has an issue (hover lists them). A collector failure collapses to null and never blocks the check.
  • Unlicensed installs. Keyless or unknown-key check-ins (previously only a log line) are upserted into the bounded stray_installs table (one row per install UUID, newest 200 kept, 90-day retention, every field length-clamped — same anti-spam posture as the error ingest). A fourth Unlicensed tab lists them (host, IP, version, Keyless vs Unknown-key badge, check-in count, first/last seen) with a 30-day count in the tab label. These are the running copies in the field that aren't Members yet.
  • Last IP. The caller's IP is stamped on members.last_seen_ip and member_installs.ip at every check-in and shown in the drill-down's Environment column beside the PHP version.

v2 batch C — remote "Update now" (signed nudge push)

Outdated installs no longer have to wait out the 6-hour poll:

  • Per-install button. In the drill-down, an outdated install that has advertised a nudge URL gets an Update now button; the Installs tab header gets Update all outdated (confirm modal, bounded to 25 sends per click).
  • How the push works. The Fleet page signs {install_id, action: 'check-updates', issued_at, nonce} with the license signing key (FleetNudge) and POSTs it to the install's _cms/update-nudge endpoint (advertised as nudge_url on every check-in). The install verifies the signature against the baked-in license public key, requires ITS OWN install id, a ±5-minute freshness window, and a single-use nonce (UpdateNudge), then spawns its normal cms:check-updates --apply in a detached CLI subprocess — never in the FPM worker.
  • Security contract. The nudge carries no code, no URLs, no version — it only means "poll your license server now," and that poll re-verifies the signed update block like any scheduled check. A forged or replayed nudge is at worst an early poll; the signature/nonce guards keep even that off the table. The receiver never touches the self-hosting updater files — it only invokes the existing artisan entry point.
  • Gating. Buttons render only when a signing key is configured (FleetNudge::canSend()), so a fleet server without the mothership's key simply doesn't get them.

v2 batch D — proactive monitoring: uptime probe, owner alerts, weekly digest

  • Uptime probe (fleet:probe, at the configurable pulse cadence via LazyCron, license server only). HEAD/GET of every install seen in the last 90 days (scheme+host from its advertised nudge URL, else https on its host; 5s timeout; any HTTP response = reachable). Bookkeeping on member_installs (last_probe_at, last_probe_ok, consecutive_probe_failures, down_since). Each pulse self-confirms: a failed probe on a site that was up triggers an immediate confirmation re-probe (after a ~1.5s pause) rather than waiting for the next pulse, so a single confirmed-failing pulse marks the install down (DOWN_AFTER = 1) — a real outage is caught within one pulse while a momentary blip that clears on the re-probe never alerts. An already-down install skips the confirmation round and just re-checks for recovery. The pulse cadence is an admin dropdown on the settings card (fleet.pulse_seconds ∈ FleetMonitor::PULSE_INTERVALS = hourly / 30 min / 15 min / 5 min / 1 min, default 30 min); a faster pulse detects sooner but only ticks as often as the site gets traffic (LazyCron is traffic-driven, not an OS cron). The drill-down shows a green/amber/red dot per install and the member row gets a red Down badge — this disambiguates stale + up = their scheduler died from stale + down = site gone.
  • The outage window (member_installs.down_since, 2026-08-25). Crossing the down line stamps when it happened and the stamp is never re-written mid-outage; the probe that finds the site answering again deliberately leaves it set, so runAlerts() can measure the downtime before closing the window. last_probe_at cannot stand in for it — that is only ever the MOST RECENT probe, which is why an outage previously had a start nobody could read (and why the drill-down's red-dot tooltip now says "down since" truthfully).
  • Owner alerts (fleet:alerts, same pulse cadence as the probe). One combined email per tick covering whatever is NEW: circuit-breaker trip (once per release), an error signature spiking across ≥3 installs in 24h (once per signature), installs crossing the 30-day stale line (once each, re-armed on return), failed update reports (id watermark), installs newly down (re-armed on recovery), and installs back up. All de-dup state lives in fleet.* Settings. A confirmed down transition also fires its alert from within fleet:probe's own run (idempotent via the de-dup), so the down email goes out the moment the outage is confirmed rather than waiting on the separately-scheduled sweep.
  • Recovery alerts (2026-08-25). A "not responding" email with no all-clear left the owner unable to tell a five-minute blip from an overnight outage — the outage that prompted this had a 4:59am alert and no second email at all. The recovered section names each install that answers again, the downtime (3h 12m), and both boundaries in site time. It fires from fleet:probe's own run too, and for the same reason as the down alert but more sharply: the figure is measured to the probe that found the site up, so a lagging alert pass would round it up by a whole pulse. Two rules keep the pair honest — a recovery is only announced for an outage that was actually reported (it must be in fleet.alerted_down_ids, so switching the down alert off doesn't produce an all-clear for an alarm nobody heard), and the window is closed outside that toggle, so it can never be left hanging and later report an outage measured in weeks. A host that becomes probe-exempt discards its window instead: it was never a real outage.
  • Weekly digest (fleet:digest). Fleet counts, outdated installs, top errors of the week.
  • Dashboard Fleet card (fleet_installs, 2026-07-24). The main dashboard shows managed installs with a count of those needing attention — behind the latest build, silent for 30+ days, or failing the uptime probe — collapsing to "All installs healthy" when quiet, linking through to this page. Backed by FleetMonitor::dashboardSummary() (three indexed aggregates plus a small version pluck, since version verdicts need version_compare) and FleetMonitor::latestKnownVersion(), which is also what this page's outdated verdicts use. See dashboard-widgets.md.
  • Settings. The Fleet Settings page (dashboard/fleet/settings, linked from the Fleet page header; fleet servers only) holds the Fleet monitoring card: alert recipients (comma-separated; blank = all admin users, same fallback as update notifications), the pulse-cadence dropdown, and per-condition toggles + the digest toggle. Saving re-registers the probe/alert LazyCron tasks at the chosen cadence immediately (upserts interval_seconds), so a change takes effect on the next tick. (This card lived on Memberships → Settings until the 2026-07 decoupling.)

All three commands no-op quietly on non-license-server installs, and the LazyCron tasks are only registered (and are actively deregistered) based on Membership::isLicenseServer().

v2 batch E — member-facing "My Installs"

Members (and agencies) get their own read-only fleet view at /members/installs (members.installs, linked from the license-key card on the member dashboard): every install checking in under their key — host, version with Update available / Up to date badge, reachability dot (from the uptime probe), an Early updates chip for canaries, and check-in freshness. A member whose team subscription owns seat members (owner_member_id) also sees each seat member's installs, grouped.

Privacy-scoped by design: no license keys, no admin notes, no error digests, and never another member's data. License-server only (the component 404s elsewhere) and skipped from the page editor's feature-page discovery like the other member-only pages.

v2 batch F — enterprise-run fleet servers (gating reworked 2026-07-24)

Fleet is no longer mothership-only. An Enterprise install enables the Fleet Management feature card (Settings → Features → Platform & Developers; requires_tier: enterprise, so the switch is locked with an "Enterprise plan" badge below that tier — guarded again in updatedEnabled and server-side in FeatureActivator). That makes Membership::isFleetServer() true (Features::enabled('fleet') && hasFleetTier(); isLicenseServer() implies it regardless of the card) — the gate the Fleet pages, sidebar item, and monitoring crons use, enforced on the routes by EnsureFleetServer.

Fleet is deliberately independent of the consumer Memberships feature. An agency can run its own site as the fleet head without ever selling memberships: the check-in endpoints ride feature:memberships,fleet (any-of), Membership::issuesLicenseKeys() (License keys toggle OR fleet server) gates key issuance/verification, and the members ADMIN routes append fleet to their any-of gate — while the PUBLIC member-account system (login/registration) stays off unless an account-bearing feature enables it.

How an agency wires it up: enable Fleet Management, create a Member per site (each gets a license key automatically — no License keys toggle needed on a fleet server), and on each client install fill the Fleet reporting fields (Settings → Updates & License): the fleet server's /api/membership/check URL + that site's key.

Then give the head a real system cron — this step is not optional. Everything the fleet does on a schedule (the probe, alerts, the weekly digest, and above all fleet:backup-scheduler, the 60-second task that admits and nudges each install) runs on LazyCron, which on a normal install is driven by page views. An agency head is typically the quietest site in its own fleet, so without a cron the scheduler ticks only when somebody happens to visit it — and a rollout stops wherever that leaves it. Add, as the web app's system user, every minute:

* * * * * <php-cli> /home/<user>/webapps/<app>/artisan lazy-cron:run

On a panel-managed host give the PHP binary as the vendor binary and only the artisan path as the command — see lazy-cron.md → "On a managed host", which also covers the three silently-failing forms (the likeliest being a path copied from another site, which leaves the panel showing a green Active badge over a job that does nothing).

Leave the trigger mode on auto. A head does not need LAZY_CRON_MODE=cron, and is better off without it: auto keeps page views and installs' check-ins as a second way to be ticked if the cron ever stops, and only auto gets the event-driven admission that runs the scheduler the moment a report frees a slot. Alerting is not the trade-off it once was — a fleet head is held to a 10-minute stale-heartbeat threshold in either mode, so a dead cron shows up on the dashboard within minutes regardless. The installs' check-in tick reduces how much a missing cron hurts, but does not replace it.

Reaching the tier. The card unlocks on tier enterprise — a paid Enterprise line item, or a comp pinned to it (mothership member edit → Comp / Exempt → Comp as → Enterprise, member_profiles.exempt_tier). A plain comp reports tier exempt, which unlocks the member features but satisfies no requires_tier gate, so the owner's own agency site could never become a fleet head until this existed (2026-09-09).

Nothing on an agency head requires the consumer Memberships feature. The probe / alerts / digest / backup-scheduler crons and the provider-held backup endpoints gate on Membership::isFleetServer() + issuesLicenseKeys() alone (they used to demand Features::enabled('memberships') and the raw License-keys toggle, which an agency never has — a head running Fleet only got a working dashboard and zero crons). memberships:prune-installs registers everywhere and gates itself on the same verdict, so an agency head's member_installs prunes like the mothership's.

The fleet channel is report-only by construction. The client mirrors its normal check-in body (host, version, install id, health, error digest, nudge URL) to the fleet URL with the fleet key as Bearer — and discards the response unread. A hostile fleet server can see what the install chose to send but can never influence licensing, updates, or local state; the mothership license channel (fixed URL + signed responses) is untouched. That is why the fleet URL/key are dashboard-editable while the license-server URL is not. The client also refuses to "fleet-report" to the same URL as its license server (pointless duplicate).

On an agency fleet server: update verdicts fall back from release.version (never set there) to the newest version the mothership has offered that install, then its own version (FleetMonitor::latestKnownVersion()). The same fallback now drives the staggered update nudge (FleetBackupScheduler::updateDueVersionFor()) and the weekly digest's outdated list, and the critical lane reads the critical flag the mothership stamped on the head's own offered update (update_critical) in place of release.critical — before 2026-09-09 all three read the raw release.* Settings and were silently dead off the mothership while the Fleet page said "N outdated". Existing anti-spam ingest bounds apply unchanged.

What an agency head does NOT do — and hides. Releases, channel soak windows, the per-install Speed override, and the release circuit breaker are SERVE-time mechanics of the update server: only the mothership offers builds, so on an agency head the Speed select is replaced by the install's own reported channel, the Rollout tab shows a single "Newest known release" tile (no Stability windows / First updaters), and Fleet Settings drops the "Circuit breaker trips" switch. The entitlement-mismatch flag ("unlicensed features") is likewise computed only on the license server — an agency's member rows are fleet identities with no plan or comp, so every managed install would otherwise read as unlicensed there.

Update health reports reach the agency head too. UpdateHealthReporter::report() mirrors each post-apply report to the fleet URL with the fleet key (the sibling of the check-in mirror), so the head's Failed updates alert fires, its Update reports panel fills, and its per-server run cap frees the update slot the moment the apply completes instead of at the deadline. The head records the report but can never trip a breaker (no release.version to guard).

Enterprise parity — the per-fleet signing key (2026-07)

An enterprise fleet server is no longer stuck report-only. Enabling the Fleet Management feature (and, self-healingly, the first Fleet-page or Fleet-Settings visit for servers enabled earlier) mints a per-fleet Ed25519 keypair (fleet.signing_key / fleet.public_key Settings — FleetNudge::ensureSigningKeypair()). Every signer now resolves through Membership::fleetSigningKey(): the mothership's env CMS_LICENSE_SIGNING_KEY first, else the minted fleet key — so signPayload(), FleetNudge pushes, and bucket-grant minting all work on an agency fleet.

The install side stays opt-in, per install. The fleet key is only honored where the install admin pinned the fleet's public key (Settings → Updates & License → Fleet reporting → Fleet public key; the key to paste shows on the fleet server's Fleet Settings page; env default CMS_FLEET_PUBLIC_KEY, which fleet-generated installers bake in so provisioned sites are pinned from first boot). With a pin, MembershipClient::pinnedFleetPublicKey() feeds exactly three verifiers:

  • reportToFleet() stops discarding the fleet response and ingests the Ed25519-verified, nonce+host-bound timing claims only: the backup_schedule block (staggered slot, destination policy, update window) and the admission run grant (the pull lane for a critical update under the per-server run cap) — the same BackupSchedule::ingest() / CmsUpdate::ingestFleetAdmission() paths the mothership channel uses.
  • The pinned head becomes the install's ONE timing authority (MembershipClient::check(), 2026-09-09): the mothership's own backup_schedule and admission blocks are dropped on every poll while a fleet is pinned. This is load-bearing, not tidiness — the fleet report fires BEFORE the license poll, and for an install the mothership doesn't manage (a plain licensed check-in never enrolls) the mothership answers an always-write managed:false, which wiped the agency's slot moments after it was applied on every single poll. Unpinned installs keep the mothership as their only timing source.
  • The pinned head can broker AI on the agency's own provider keys (2026-09-09). The head already answers the AI proxy (api/ai/{dialect}/v1/{path} rides feature:memberships,fleet + issuesLicenseKeys()) and signs a managed_ai block into its check-in response from its own Managed AI settings (managed_ai.enabled, managed_ai.{dialect}_key, models, comped credit, per-member ai_credit_cents / ai_images_all_users on ITS member rows). With the head pinned, ingestFleetTimingClaims() ingests that block with ManagedAiClient::SOURCE_FLEET, MembershipClient::check() skips the mothership's block while the source is the fleet, and every brokered lane presents the FLEET key (ManagedAiClient::credential()) — the mothership license key is never sent to a fleet URL, so the pin directs only a credential the head already issued. A head that withdraws its block (or an admin un-pinning) hands the lanes back to the mothership on that same poll; an unreachable head keeps the last state. Metering lands on the head's ai_usage_daily under the fleet identity, so the agency pays and meters its clients; the install's meter shows the head's allowance. The Managed AI settings card still lives on Memberships → Settings (gated feature:memberships) — a Fleet-only head configures it via Settings until that card moves.
  • The mothership knows. Every license poll carries fleet_host + fleet_pinned; the ingest stores the pinned host on member_installs.fleet_host (always-write; null when unpinned, malformed, or when the pinned host is this server itself — a fleet head sees its own installs' mirrored check-ins). A pinned-elsewhere install is skipped by managedInstalls(), answered managed:false by claimBlockFor(), never admitted by the run cap, excluded from the Not scheduled tile, and its drill-down shows a Managed by webprojoe.com badge in place of the Scheduled/Unscheduled toggle (a stale toggle click is refused server-side). So moving an install to an agency head needs no manual un-manage on the mothership: the install's next poll hands it over. (2026-09-09 — the migration of the owner's own fleet was the trigger.)
  • The signed-nudge receivers (_cms/update-nudge, _cms/backup-nudge) accept the pinned key beside the mothership keys — a fleet's check-updates nudge still only means "poll the mothership now."
  • FleetGrant accepts bucket-upload grants minted by the pinned fleet.

What the pin can never grant: license verdicts, update blocks, and release zips verify exclusively against the baked-in mothership keys (trustedPublicKeys() / ReleaseSignature) — the pinned key is deliberately excluded, so even a fully-trusted fleet server cannot forge activation or deliver code. Unpinned installs behave exactly as before: the fleet channel is report-only and the fleet's Rollout tab notes that its schedule only reaches pinned installs.

When the agency's plan lapses (2026-09-09)

Traced end to end, because "does a client install know it's the agency's key that lapsed?" is the natural question. It doesn't need to — it holds the agency's key. On the agency's next poll the mothership answers active:false, tier:null, licensed:true for that key, and every install using it records membership.is_member = false on its own next poll (within 6 hours; the beacon only accelerates release discovery, not licensing).

  • Every install, the head included, drops to the free level. The plan gate (Features::premiumLocked()) turns off every opt-in feature — AI chat, the assistant, the SEO scanner, e-commerce, the members area — while default-on features and the site itself keep working. The owners' configured switches are untouched; paying again flips them back on at the next poll. Updates keep flowing (licensed stays true; a recognised key always gets builds).
  • The head stops being a fleet head. tier is no longer enterprise, so isFleetServer() is false: the Fleet page 404s, the probe/alerts/digest/scheduler crons no-op, and the head's check-in and AI-proxy endpoints 404 (issuesLicenseKeys() false).
  • Managed installs degrade, never strand. Their fleet reports fail silently, so they keep the last signed timing state: backups still fire at the persisted slot from each install's own heartbeat catch-up, and an unattended update waits for an admission grant the head can no longer issue until the 24-hour deferral cap restores autonomy. Fleet-brokered AI is already off with the features; the mothership's own AI block is absent too (the key is inactive), so nothing is left pointing anywhere usable. Un-pinning the head on an install hands timing back to the mothership immediately.
  • Renewal reverses all of it on the next polls: the head's tier returns, its crons and endpoints come back, and the installs' features and fleet-brokered AI resume with no re-setup.

A head that goes dark while the agency still pays is the other half, closed the same day: the install stamps membership.fleet_last_ok_at on every VERIFIED fleet answer (a 200 without a valid signature does not count), and once that stamp is older than MembershipClient::FLEET_DARK_AFTER_HOURS (24 h — four missed polls) the head is treated as dark and the install goes local: it clears the fleet's timing state and runs its own automated backup schedule, and applies updates on its own with no admission to wait for. It never takes a mothership backup slot or admission grant — dark head or not (pinnedFleetConfigured() drops those blocks unconditionally): the mothership doesn't know which server a fleet's installs share, so it is never the right scheduler for one. AI falls back to the mothership's proxy only if the mothership already counts the install's key as a paying subscriber (managed AI is offered to active keys alone, metered against that member's allowance) — which for most fleet installs it will not, so they simply lose AI until the head answers again or the owner adds their own provider key under Settings → API Keys. The head's first verified answer restores everything on that same poll. The Updates page shows the last contact and says when the site is on its own. Locked by FleetDarkHeadTest.

SEO & traffic telemetry (2026-07)

Two more optional summary blocks ride every license check-in (and therefore mirror to agency fleet servers automatically, since the fleet report is the same body):

  • seo — the install's latest completed SEO scan reduced to {score, errors, warnings, scanned_at} (SeoTelemetry). Null (block omitted) when the SEO Scanner feature is off or no scan has completed.
  • traffic — rolling {views_7d, visitors_7d, views_30d, visitors_30d} from the analytics rollups (TrafficTelemetry). Four aggregate ints only — never paths, referrers, or visitor-level data. Because traffic is the client's business data it rides an owner-facing toggle (Share traffic summary, telemetry.share_traffic, default on, beside the error-sharing switch in Memberships → Settings → Membership client).

Ingest whitelists and clamps both blocks onto scalar member_installs columns (score 0–100, counts capped, future scan stamps clamped to now) so the Installs tab can sort and filter in SQL. An absent block nulls its columns — an install that disables the scanner or the traffic toggle stops displaying stale numbers on its very next poll.

On the Fleet page:

  • Installs tab gains a Traffic (30d) column (member row = sum across its installs) with a click-to-sort header toggling between check-in recency (default) and 30-day views — the "client sites ranked by traffic" view. Persisted via ?sort=traffic.
  • An eighth SEO critical stat tile counts members with an install whose latest scan landed in the critical band (score < 50), and filters the table like the other tiles.
  • The drill-down shows each install's banded SEO score badge (green ≥90 / amber ≥50 / red <50; hover = error/warning counts + scan age) and its 30-day views · visitors.

Accessibility telemetry & client report routing (2026-09-09)

Two additions that make the weekly digest an agency's whole client-health picture rather than a server-health one.

  • accessibility — a third optional summary block on every check-in, the exact twin of seo: the install's latest stored report reduced to {score, issues, pages_scanned, scanned_at} (AccessibilityTelemetry). Reads the newest report of EITHER trigger, because a manual "Generate Report" is the more current picture when someone has just run one. Ingest clamps it onto member_installs.a11y_score / a11y_issues / a11y_pages / a11y_scanned_at, and an absent block nulls them like the seo/traffic columns. The drill-down shows a banded A11y badge beside the SEO one.

  • The weekly digest gained three ranked tables — SEO scores and Accessibility scores (both lowest first: the digest is a worklist, not a leaderboard) and Traffic (last 30 days). Until this landed, SEO and traffic existed only on the Fleet page, so the digest could tell an agency which install was burning CPU but not which client site had lost its ranking.

Who receives a client site's own report emails is now the fleet's decision, made when the site is added to it: FleetReportPolicy (head side) resolves member_installs.report_policy else the fleet-wide fleet.default_report_policy, and signs a reports claim — {policy, agency_email} — into the check response. ReportRecipients (install side) ingests it from the VERIFIED claims only and is consulted by all three owner-facing senders: the Analytics scheduled website report, the SEO score-regression alert, and the new weekly accessibility report.

Rules that are load-bearing:

  • Signed, like the backup slot. The block redirects a client's outgoing mail, so an unsigned one would let a relay point a website report anywhere.
  • Only a MANAGED install is told anything (backup_managed, and not pinned to another head). Routing the mail of a site this fleet doesn't run would be redirecting post nobody handed us.
  • Client-only sends no block at all — that is what an install already does on its own, so saying it would require an agency address for a policy that never uses one.
  • A policy that resolves to nobody degrades to the client. An agency policy with an empty address list would otherwise compose the report and send it nowhere, raising no error and emptying no inbox anyone would miss.
  • The install's OWN recipients are never re-validated — ReportRecipients::normalize() trims and de-duplicates but deliberately does not run FILTER_VALIDATE_EMAIL, which rejects root@localhost, the default super user on a fresh install. Routing a report must never be able to remove a recipient that was already receiving it. (Caught by SeoScannerAlertTest during the build.)
  • The fleet asking for reports IS the opt-in for the accessibility email (AccessibilityReportMailer): requiring the client's own switch as well would mean an agency's choice silently did nothing on every site it manages. It links to the DASHBOARD report, never a minted share link — a share token exempts its report from pruneOldReports() forever.

Locked by tests/Feature/FleetClientReportsTest.php.

Fleet-scheduled SEO scans. Managed installs (the same opt-in backup_managed set) also get a weekly SEO-scan slot riding the signed backup_schedule claim (seo_day/seo_time): by default the backup minute three nights after the backup slot, overridable per install via the day picker under the drill-down's SEO badge (Auto / Sun–Sat, stored on member_installs.seo_slot_dow). On the install, BackupSchedule::fleetSeoSlot() (written only from the signed claim; the admin's self-managed opt-out wins) feeds the scanner's scheduled-scan gate — the scan fires once per weekly slot occurrence in the site's own timezone instead of "a week since the last scan". The install's own seo_scanner.scheduled_enabled toggle still wins: the fleet moves the clock, it never forces scanning on. No nudge is needed — the scanner's own 300-second LazyCron checks the slot locally.

Backup scheduling — fleet-orchestrated, staggered

The fleet head schedules and staggers every managed install across the fleet's three cadences: a daily local-snapshot slot (which doubles as the update-apply moment — the same slot staggers the pre-update snapshot each update triggers), and a separate weekly off-site upload slot. Server side is FleetBackupScheduler; install side is documented in backups-and-revisions.md → "Fleet-managed timing".

  • Which installs. Scheduling is opt-in per install (backup_managed, default off) — the fleet never annexes every install that merely licenses against the server. An install auto-enrolls in exactly one way: its check-in arrived over the fleet-report channel (fleet_report: true on the mirrored body — the install admin's explicit "this server manages me" act). Anything else — including every standalone-plan install licensing against the mothership — stays unmanaged until the fleet owner flips it on in the Installs drill-down. The owner's toggle is never overwritten by later check-ins (create-only enrollment), and unmanaged installs still stagger on their own: each install's LOCAL default slot is a deterministic pseudo-random night derived from its install id (see backups-and-revisions.md → "Scheduled backups").
  • What it computes. Two slots per managed install, each packing a different resource: the backup slot — a minute inside the 00:00–05:00 local maintenance window, packed on a single-day absolute (UTC) ring per server because every install fires every night (day-of-week is no longer a packing dimension; backup_slot_dow is still assigned for clients predating the daily cadence, which read the slot as weekly day+time). One ring per effective server key, with singletons and unidentified installs sharing one pool ring — see "Multi-server fleets" for the pool rule and why it exists. Then the off-site slot — a day-of-week + minute anywhere in the 24-hour day, packed on the full-week ring quiet-hours-first (nearest 03:00 local, spilling toward midday only as the night fills; days fill fullest-first so uploads cluster before opening another day), because an upload never takes the site offline and must not contend for the maintenance window. Slots are spaced by each install's real measured backup duration (reported on the license poll from the second — first incremental — snapshot onward; a 5-minute default until then). Placement uses escalating lanes, never refusal: everyone gets a slot alone if one exists, else stacked 2-wide, then 3-wide — max concurrency is only the depth at which the capacity warning trips, not a cap. Placement is stable (a still-valid assignment is never moved) and a daily compaction sweep scoots later backups earlier — only ever into empty runs — as conservative default widths are replaced by measured durations.
  • How it delivers + fires. Both slots ride back to the install inside the Ed25519-signed membership response (a backup_schedule claim carrying frequency: daily, the slot day/time, and offsite_day/offsite_time). The fleet:backup-scheduler LazyCron task (fleet server only, ~every minute) assigns slots and then nudges each install at its DAILY slot (within slot + 90 s grace, slot + 6 h) — outside that window the work waits for the next night rather than firing the whole fleet at once when a late fleet head returns; once per period, ≤5 send attempts/tick — failures count too, so unreachable hosts can't hold the tick in HTTP timeouts) via a signed run-backup push to _cms/backup-nudge — the backup-nudge sibling of the update nudge. A failed send puts the install on a cache-backed exponential backoff (5 min doubling to 6 h; cleared on success) instead of retrying every tick. If the fleet head can't reach an install, the install's own heartbeat catch-up fires the same backup at the same persisted slot; a period-scoped claim lock means the two paths never double-back-up. The off-site slot gets no nudge: the install's own backups:offsite tick enqueues the newest snapshot when the slot arrives (a quiet, traffic-less site falls back to riding its daily backup nudge — CreateBackupJob pushes off-site whenever the weekly moment is due).
  • Update timing. A pre-update snapshot is a backup, so update-applies ride the same staggered slot: at the slot the fleet head nudges any outdated install to apply now (a check-updates nudge — preferred over the backup nudge that tick, de-duped per night via update_nudged_period), so the pre-update snapshots stagger across the fleet exactly like the weekly backups rather than all firing when a release drops. Only installs that report auto-install ON are nudged, and the nudge carries mode: scheduled in its signed claims so the receiver runs cms:check-updates --apply-scheduled — the operator's auto-install toggle and the minor/patch-only gate are enforced on the install; only the window is bypassed (the nudge lands at the slot). A manual-mode install is never remote-applied. The fleet-assigned cms.fleet_auto_install_hour (same signed block) is the mothership-down fallback window. Critical updates still bypass the window — but not the per-server run cap below; the manual "Update all outdated" button remains a deliberate go-now --apply override; and a successful update marks the day's backup satisfied (the install's BackupService::createSnapshot calls BackupSchedule::markRan, and the nudger stamps backup_nudged_period alongside update_nudged_period) so the two never double up.
  • UI. The Rollout tab has a Backup Scheduling card — master enable, default backup minutes, safety gap, night-window start/end, the stacking alert threshold, the Max concurrent runs per server select (the run cap below), and a window-filling-up warning (fleet.capacity_warning — informational only; lanes escalate, nothing is refused). A Live runs strip shows each server group's in-flight count vs the cap, critical installs waiting in line, and the stuck-run reclaim counter. The Installs tab drill-down shows each install's assigned slot + measured duration, a per-install Scheduled / Unscheduled toggle (backup_managed), and an Update/Backup starting/running badge while the install holds a run slot; an install whose admin opted out shows self-managed.
  • Respecting the admin. The mothership never overrides an install whose admin flipped "Manage timing myself" (backups.self_managed, reported on the poll) — those are excluded from scheduling and never nudged. Turning an install off in the drill-down (backup_managed false) hands it back to its own local schedule on the next poll.
  • Gating. No-ops instantly off the fleet server (mirrors the probe). Slot delivery needs the license signing key twice over — the slot claim only rides the SIGNED membership response, and the nudge is a signed push — so a fleet server without one (an agency's report-only fleet server, which by design cannot influence its installs) no-ops entirely: no slots are assigned, and the Rollout tab shows a callout explaining that installs fall back to their own staggered local defaults. This keeps the dashboard from presenting dead bookkeeping as live scheduling.

Run concurrency — capping simultaneous updates/backups per server

Slot staggering spreads scheduled starts, but nothing above stops many heavy runs from being in flight at once on one box — a critical release (which skips the nighttime window) or bunched slots could put dozens of snapshot-then-extract-then-migrate runs on a single server simultaneously. The per-server run cap closes that: the owner sets X = Max concurrent runs per server on the Rollout card (fleet.max_concurrent_runs, default 2, Off restores the uncapped behavior), and the fleet head only starts new unattended runs on a server while in-flight < X. Server side is [FleetRunCoordinator.

  • Only automatic work is capped. A person clicking Update Now (locally or on the fleet page), cms:check-updates --apply, or a --force backup is never delayed — the cap governs the scheduled + critical auto paths exclusively.
  • Grouping by physical server. Installs group by their effective key (a manager override wins): member_installs.server_override when the fleet manager set one in the Installs drill-down, else the machine fingerprint each install reports on its license poll (server_id — a hash of /etc/machine-id or the hostname, so every install on one box reports the same value; proxy- and NAT-independent), falling back to the CDN-aware observed check-in IP (CF-Connecting-IP when the fleet server sits behind Cloudflare, else the connection IP) for clients predating the field. Installs with no resolvable identity pool into one shared "unknown" group, capped like any other — never uncapped. A single server may also narrow or widen X for itself via fleet_servers.max_concurrent_runs — see "Multi-server fleets".
  • The installs are the head's heartbeat (2026-09-09). The scheduler tick is a LazyCron task, and LazyCron runs after page views — which a quiet agency head barely gets. Watching the first critical drain through webprojoe: four nudges at 19:43, then nothing until the mothership's half-hourly probe woke the head at 20:00, then three more. So the two API endpoints the fleet actually hits — the check-in (MembershipCheckController) and the update report (UpdateHealthReportController) — now queue the same deferred LazyCron::deferRunIfDue() pass a page view does, on a fleet head only. A nudged install reporting back is what admits the next one, so the drain largely paces itself on the cap rather than on traffic. Locked by FleetHeadCheckInTickTest. It narrows the window; it does not remove the need for a cron on the head. LazyCron::gate() caches its due-count for 60 s and a finished run clears it, so right after a batch the gate repopulates with due_count = 0 — the scheduler just ran and is not due for another minute. Check-ins landing in that window defer nothing, and if the batch that just finished was the one whose reports would have woken the head, the chain ends there: everyone it nudged has reported, and nobody else is due to poll for hours. That is exactly how the 1.5.15 critical drain stopped one install short on webprojoe (2026-09-10) — five clean pairs at the cap, then silence, with the eleventh install critical-queued and every gate green. Full trace in lazy-cron.md.
  • Admission, not just sending. The tick reserves a slot (run_state/run_kind/ run_claimed_at/run_deadline_at on member_installs) before nudging, so the next tick can't over-admit while an install is still spinning up. The nudge carries an admitted claim inside its signature; the install turns it into a local time-boxed grant that its own auto gates check (cms-updates.md → run admission, backups-and-revisions.md → fleet-managed timing). Priority when a server has free slots: critical updates first (FIFO by first-seen), then scheduled updates by slot time, then backups — the per-install "update precedes backup" rule, generalized per server.
  • Critical gets in line now — not tonight, and not past the cap. A critical release enters the queue immediately regardless of the nighttime window. Push lane: the tick nudges queued installs as slots free. Pull lane: an install polling with a due critical is admitted at poll time when its server has room — the grant rides the signed claims, covering installs the tick can't push to. Either way X still binds.
  • Slots free themselves. Update slots release in real time when the §D.4 health report arrives (a complete report also records the install's new version on the fleet row, so due-ness clears without waiting for the next poll; a failure starts a 30-min apply backoff). Backups have no completion report, so their deadline — twice the install's measured backup width — is the release mechanism. Every reservation carries an absolute deadline that the occupancy count enforces by itself: a crashed run's slot stops counting the moment it expires, even if no tick ever sweeps it, so a stalled (e.g. traffic-ticked, cron-less) fleet head can only under-admit, never wedge. The tick's reclaim sweep is hygiene + observability — and it grades what it sweeps: an update reaching its deadline went silent mid-flight (log warning + the fleet.run_reclaims_total counter on the Live-runs strip), while a backup reaching its deadline is simply how a backup slot is released, so it logs at info and is never counted. Warning on those put a nightly "stuck backup slot" in every fleet head's log and filled the counter with routine work, which is how a real stuck run would hide.
  • Poll evidence is freshness-checked. The install's poll reports update_status / backup_status; a status may confirm or early-free a reservation only when its claim stamp (update_started_at — the install's own claimRun() stamp) postdates the fleet's reservation, so the pre-apply poll (still carrying the previous run's status) can never free a slot that was just granted.
  • The install is never stranded. All install-side waiting is capped: an update awaited past 24 h applies autonomously (the same deferral-cap shape as the install window), and a backup awaited past the fleet's own 6 h nudge window runs locally — a dark mothership degrades to today's behavior, never to "no updates, no backups."
  • Back-compat. Clients predating this release ignore admission entirely (they group by IP and self-apply as before), so the cap becomes fully binding as the fleet updates onto the release that ships it — no flag day. Turning the cap Off (or an old mothership sending no admission_managed key) releases every install on its next poll: the flag is always-written from the signed schedule block.

Multi-server fleets (2026-09-09)

A fleet manager with 100 installs across 4 boxes used to be staggered as if all 100 shared one: the run cap already grouped by server, but the slot packer ran one occupancy ring for the whole fleet, so slots stacked, the capacity warning tripped, and updates landed later in the night than they needed to. Nothing broke — it was over-conservative. This section is what replaced it.

  • The effective server key. Everything that groups by server — the run cap, the packer, the dashboard — reads MemberInstall::effectiveServerKey(): the manager's member_installs.server_override when set, else the install-reported server_key (the machine fingerprint, else the CDN-aware observed check-in IP), else the shared unknown group. Its SQL twin is the onServer($key) scope, and the UNKNOWN branch means no override AND no reported key — so an override is cleared with '', never with the literal string 'unknown', which would group the install nowhere.
  • Servers are rows. fleet_servers carries a human label, an optional per-server run cap, and first/last seen. Rows auto-register at check-in (FleetServer::touchSeen(), at most one write an hour per key), so the table fills itself with the keys the fleet reports. The manager can also create one by hand — key manual:{slug} — which is how a fleet whose reported identity is fragmented gets regrouped onto the box it actually shares. FleetServer::pruneStale() (run by the daily memberships:prune-installs) drops a row only when no install references it by effective key and it has not been seen for 90 days, so a hand-made grouping survives regardless of when it was last "seen". A per-server cap is clamped 1..8 (FleetRunCoordinator::capFor()): a server can re-size the cap but never switch it off — that stays the fleet-wide Off.
  • Packing is per RING, with a safety pool — and this is the load-bearing rule. Per-server packing makes server identity load-bearing, and the reported identity can be wrong in the fragmenting direction: a fleet of per-site containers, each reporting its own machine id, would give every install a ring of its own. Every one of those rings is empty, so every install lands on the window's first minute and the whole fleet fires together — and the per-"server" run cap would not hold them back either, because each of those groups is one install. So FleetServerGroups::rings() gives an effective-key group of two or more installs its own ring and puts singletons and the unknown group into one shared pool ring. A fragmented fleet therefore degrades to exactly the pre-2026-09 behaviour (one ring for everyone), and a genuine one-site-per-VPS fleet is merely over-staggered — the safe direction to be wrong in. The run cap keeps its per-effective-key semantics (it never pooled singletons, and changing that would regress legitimate fleets).
  • Off-site upload slots stay fleet-wide. An upload never takes a site offline, and when the destination is the head's own disk or bucket the contended resource is the sink, not the source box — so partitioning those by source server would optimise the wrong thing.
  • Warnings, never refusals. The capacity warning is per-ring: fleet.capacity_warning stays the bool, and fleet.capacity_warning_rings names the rings that stacked past the depth, which the Rollout card appends as "Tight: …". Placement still escalates lanes rather than refusing anyone.
  • UI. Rollout tab → a Servers table under the Live-runs strip: label (editable), key (manual badge for hand-made ones), install count by effective key, the per-server cap (Inherit (X) or 1/2/3/4/6/8), and last seen — plus Save servers. The Live-runs strip shows each group against its own cap. Installs tab → a server filter beside the search (?server=, counts by effective key, plus Unknown server), and in the drill-down a Server select: Reported (…), one option per server row, and New server… which reveals a name box and creates the manual: row + assigns it in one go. An install carrying an override wears a small overridden badge whose title names the key it actually reports.
  • Fragmentation callout. When at least 3 managed installs — and at least half of them — are each alone on their reported server, an amber callout on the Rollout card asks the manager to group the ones that really do share hardware, with a Dismiss that stores the count in fleet.fragmentation_ack; it reappears only if the count grows.

Adding a website — the preconfigured installer

The Installs tab's Add a website button generates a personalized copy of the WordPress installer plugin (FleetInstallerController, fleet servers only): set up WordPress on the new host, upload the zip, click Install — the site activates, reports to this fleet, and enrolls in managed backup scheduling automatically, with nothing to type.

  • What gets baked. The plugin's PRECONFIG placeholder is replaced with base64 JSON. On the license server (the mothership): the picked member's license key + the fleet's enroll token. On an enterprise (agency) fleet server: the agency's own mothership key (cms.license_key — their plan covers their client sites) for the license channel, plus this server's check endpoint + the picked member's key as the fleet-reporting pair. A stock (non-generated) plugin still shows the placeholder → behaves exactly as before.
  • How it flows. The plugin pre-verifies the baked key against the mothership, shows a "Preconfigured by your provider" note instead of the license field, and carries the fleet fields through wpcms-prefill.json → install.php → CMS_FLEET_API_URL / CMS_FLEET_KEY / CMS_FLEET_ENROLL_TOKEN in the new site's .env. The dashboard's Fleet reporting fields (Settings → Updates & License) override those env defaults whenever an admin edits them — the same Setting-or-config pattern as the license key.
  • The enroll token. A per-fleet secret minted once (fleet.enroll_token Setting) and baked into every installer this fleet generates. An install whose check-in carries the matching token auto-enrolls in backup scheduling on first sight (backup_managed, create-only — the owner's drill-down opt-out always wins). It proves the fleet OWNER provisioned the install, which is what keeps enrollment opt-in on the license channel: a plain licensed check-in without the token (or a fleet report) never enrolls anyone.

Backup destinations — provider-held copies

By default each site keeps its snapshots on its own hosting, written by its own unix user — which means ransomware running as that user can shred the site and its backups. The Rollout tab's Backup Scheduling card lets the fleet owner add a provider-held copy per managed install (the same opt-in backup_managed set as scheduling):

  • Keep backups on each install (default) — no provider copy; the site's own Backups page off-site card (its own S3/Dropbox/FTP) remains available as ever.
  • Also copy to this server's disk (fleet_disk) — after each snapshot the install pushes its bundle zip to the fleet server in 8 MB chunks (api/fleet/backup-offset + api/fleet/backup-chunk, license-key Bearer, append-only .part + finalize-by-rename, size caps, disk-free guard). Stored under storage/app/private/fleet-backups/{install_id}/ (CMS_FLEET_BACKUPS_PATH overrides), owned by the fleet manager's user — the compromised site cannot touch it. Retention is keep-N per install, pruned on finalize.
  • Also copy to shared object storage (fleet_bucket) — ONE bucket for every site the fleet owns, so adding a website needs zero storage setup. The install never holds the bucket credentials: it asks api/fleet/backup-grant for an Ed25519-signed presigned PUT scoped to its own prefix ({prefix}/{install_id}/…, 30-minute expiry, nonce-bound — see FleetGrant) and streams the zip straight to the bucket. Write-only by construction; the fleet server prunes each install's prefix once the object has landed (api/fleet/backup-confirm), and reconciles at the next grant for installs that never confirm.

Retention integrity. Both modes treat the install as potentially compromised, because that is the case the feature exists for. Nothing the install sends decides which copies survive: retention orders on the timestamp the FLEET SERVER observed (disk mtime / bucket LastModified), never on the caller-supplied filename — a lexicographic sort would let a planted webprocms-backup-9999-12-31-…zip outrank every genuine copy and survive every prune. Bundle names dated in the future are refused at intake, the prune never runs speculatively before an upload (so an abandoned upload cannot retire a real copy), and grants are capped at FleetBackupDestinationController::GRANTS_PER_DAY per install per rolling day. Residual: a compromised install can still push junk bundles at its allotted rate and churn a keep-N window forward — bucket versioning or an object-lock policy is the answer where that matters.

Delivery + trust. The policy (mode + API base URL) rides the same SIGNED backup_schedule claim as the slot — a relay can never redirect a site's backups to a rogue store, and an unverifiable grant fails the upload loudly instead of exfiltrating a copy. On the install the two destinations are ordinary OffsiteDestinations (FleetDiskDestination / FleetBucketDestination) that enable themselves only under a delivered policy, so the bundle/queue/resume/retry machinery is the existing off-site pipeline. The install side is deliberately write-only — listing, download, and deletion of provider copies happen on the fleet server only.

Restore path. A provider copy is a standard bundle zip: download it from the fleet server's disk (or the bucket) and feed it to the site's Backups → Restore from Uploaded Backup.

Browse panel (v2, 2026-07). The Installs drill-down's archive-box button opens each install's Provider-held backups panel: every stored copy from both locations (fleet disk + bucket, newest first) with size and age, a header rollup (count · total · newest, plus an amber Stale badge when the newest copy is older than two weekly cycles — the silent-push-failure detector), per-copy Download (disk copies stream from the fleet server; bucket copies redirect to a 5-minute presigned GET, so credentials never leave the server), and per-copy Delete behind a confirm modal (owner-side only — the install still can never list or delete its copies). Helpers live on FleetBackupStorage (copiesFor / bucketDownloadUrl / deleteCopy / bucketTotalBytes, sharing the promoted bucketDisk() builder with the grant path), and the Rollout card's bucket mode now shows the prefix's total usage beside the disk mode's existing line.

Resource usage — per-install CPU, DB time, and disk (2026-08)

The question this answers: on a shared box, which install is burning the server's CPU and filling its disk. Load average and CPU % are server-wide — every install on one machine reports the same number — so the per-install differentiator is ResourceMeter: a getrusage() baseline taken at provider registration and committed by a shutdown function (after deferred work, so LazyCron's background passes are counted), giving each process's exact user+system CPU delta with no sampling, no shell-outs, and no /proc access. Database wall-time rides the same commit from the framework's always-on Connection::totalQueryDuration() — the closest an unprivileged install gets to per-install MySQL load (mysqld's own CPU is unattributable without root; a lock-bound query is load even when it burns no CPU).

Client side. Commits land as cache increments in per-day buckets (site-timezone days). The minutely resources:sample LazyCron tick samples the 1-minute load average (server-wide context, grouped by server_key on the fleet head) and folds finished days into the resources.daily Setting ledger (8 days). The daily resources:measure-disk tick weighs the install's whole tree — du -sk when the host allows subprocesses, a budgeted PHP walk otherwise — the per-install number the whole-filesystem disk_free_pct health signal can't give. ResourceTelemetry ships the last 3 finished days + today's partial + disk/cores/load as the check-in's resources block, gated by the telemetry.share_resources toggle (Memberships → Settings, beside the error/traffic toggles; default on).

Mothership side. MembershipCheckController::recordResourceDays() upserts the day entries into install_resource_daily (the ai_usage_daily pattern: one row per member/install/day, day deliberately uncast) — monotonically per counter, so a smaller replay or a mid-day cache clear on the install never shrinks a recorded total — and promotes disk_bytes / cores / load_1m to member_installs columns (nulled when the block is absent, like seo/traffic). Bounded at ingest: ≤4 day entries per check-in, each within a week of the server's clock, every number clamped. History is pruned past InstallResourceDaily::RETENTION_DAYS (90) by memberships:prune-installs.

Where it surfaces. The Fleet dashboard's Resources tab: per-server cards (latest load, cores, install count, summed disk, keyed by the run cap's server_key) and a table of every reporting install ranked by 7-day CPU, with CPU-yesterday, DB wait, requests, and disk. The weekly digest gains three ranked tables: Busiest installs (per-day CPU/DB/requests averaged over reported finished days), Largest installs (disk), and Managed AI usage (7-day requests + tokens per member/host, heaviest first — renders only when ai_usage_daily has rows, so non-AI fleets never see it).

What it deliberately is not: a monitoring agent. Metering is one cache increment per request; losing the cache costs at most one partial day; and CPU covers PHP only — attribution of mysqld's own time is out of scope by design.

CPU-hog alert (2026-08). A seventh condition in FleetMonitor::runAlerts() (fleet.alert_cpu_hog switch on Fleet Settings, default on): one install consuming ≥ CPU_HOG_SHARE (60%) of its server group's total measured CPU for a finished day, with an absolute floor (CPU_HOG_FLOOR_MS, 1 CPU-hour) so a quiet server's biggest-of-tiny-sites never pages, and a ≥2-reporting-installs group requirement (a lone install is trivially 100%). The pooled "unknown server" group is skipped — shares of a group that mixes unrelated machines mean nothing. Relative-not-absolute is the point: the same daily burn is normal on a busy box and alarming on a quiet one. De-duped per (install, day), so a chronic hog re-alerts each day it stays on top until dealt with.