Processing...

 Oversight: Modern Infrastructure Monitoring from GEN

Every alert earned, every outage explained

Oversight: Every Alert Earned, Every Outage Explained

Oversight is infrastructure monitoring from GEN that only raises the alarm when the evidence agrees, and when it does, sends one message that explains the whole incident. It watches servers, clusters, storage, networks and services from as many places as you choose, and it is managed end to end by GEN.


We built Oversight from scratch because the monitoring we had been using for years kept failing in the same ways: false alarms from one flaky link, alert storms that buried the cause, green dashboards sitting over probes that had quietly died, a mistyped threshold paging an engineer at three in the morning, and licences that charged more every time the estate grew.


It is what we use in our own datacentres. The Proxmox clusters, Ceph storage, mail platforms and network equipment that carry our customers' services are watched by Oversight, so every feature described here has been proven on infrastructure we are directly answerable for.


The Modern Way to Monitor

Most monitoring was designed for a world of single servers and fixed ports: one poller checks a thing, and if the check fails, somebody is told. Today's estates are clusters, replicated storage, virtual guests that migrate between hosts, and vendor APIs that describe their own health in detail. Asking a single poller whether a port answers tells you very little about any of that.


Oversight is designed for estates as they are now. It reads what your platforms say about themselves, asks from several vantage points before it decides anything is wrong, understands that a guest in backup is not a guest that has failed, and treats the absence of data as exactly that rather than as good news. The result is monitoring your team keeps reading, because it does not cry wolf.


Alerts You Can Believe

One probe that cannot reach a server tells you about the probe. Oversight attaches each check to probe groups in different places, forms a separate view from each, and decides the state by how many of them agree:


Failing viewsDefault stateWhat it means
Exactly oneOKOne vantage point struggling is the internet, not an outage
More than one, not allWarningSomething real is happening
Every reporting groupCriticalNobody can reach it
  • The same rule at every level. Sensors roll up into devices, devices into groups, and groups into sites, each with its own mapping. A server goes red on any failed check; a resilient switch fabric shrugs off one lost path; twenty web servers behind a load balancer only matter when several fail.
  • Silence is never read as health. A probe group that did not report is left out of the count, never counted as healthy.
  • Unknown is never downtime. A sensor we cannot hear from is shown as no data, never drives its parent into alarm, and never counts against availability.
  • A typo does not wake anyone. Configuration is validated on save, and anything that still cannot run is suspended after a single attempt and flagged, not paged.
  • The worst rule wins. Every rule is tested, so the order they were written in can never hide a critical behind a warning.
  • Fix a rule, fix the history. Raw readings are kept and rules applied centrally, so a corrected threshold can be re-evaluated against past data.

One Incident, One Message

When a check fails, its device, group and site can all follow. Most monitoring sends four alerts, or "Sensor X is down, and 45 others". Oversight folds the incident into one message per rule, drawn as a tree: the site, the group, the device that went critical and the check beneath it that took it there, worst first.


Primary DC                        WARN
└ Proxmox                         CRIT  2 devices
  ├ pve-02                        CRIT  4m
  │ ├ Cluster quorum    offline 1  CRIT
  │ └ Guests  133 proxy stopped    CRIT
  └ pve-04                        WARN  1m
    └ Ceph health     HEALTH_WARN  WARN
Offices                             OK
└ fw-01  Ping                        cleared

  • Delivered by email over your own SMTP relay, to Matrix and Rocket.Chat, or as a webhook to anything with an API, such as an n8n flow.
  • Charts and the poll log from the fifteen minutes before the alarm travel with the message.
  • Escalation is a second rule on a worse state, to different people by a different channel.
  • Recovery goes only to the people who were told about the failure, and repeats are bounded.
  • Office hours and maintenance windows run on local time, correctly through the clock change.
  • Any rule can be sent as a test to its real recipients before a real incident tests it for you.

Built for the Estate You Actually Run

Oversight understands clusters, storage, backups and telephony, not just ports. One fetch fans out into as many separate alarms as the answer deserves, so five sensors reading five fields of one document cost one request.


  • Proxmox VE clusters: quorum, and online and offline nodes by name
  • Proxmox guests: every VM and container with its own rule, with backup, migration and snapshot recognised as expected states
  • Proxmox Backup Server, watched alongside the cluster it protects
  • Ceph: health, monitor quorum, OSDs up and in, and placement groups
  • Asterisk: every endpoint and trunk, with the names of those down in the alert
  • Any REST or SOAP API, including Redfish, HPE iLO and Dell iDRAC
  • SNMP v1, v2c and v3, choosing objects from vendor MIBs rather than typing OIDs
  • MySQL, MariaDB and MongoDB queries, such as replication lag or queue depth
  • FTP and FTPS with file age, so an eleven-day-old backup is a failure
  • HTTP(S), DNS, SIP, RDP (including NLA enforcement), SMTP, IMAP4, TCP and ping
  • Email delivery, sending a real message and following it to the mailbox, or all the way back
  • Legacy equipment, with older TLS allowed per sensor as a deliberate choice

The estate wheel shows every site, group and device at a glance, coloured by state, with your own logo at the centre, and availability is reported alongside how much of the time we could actually see it.


Secure by Construction

  • Outbound HTTPS only. Probes connect out; nothing connects in, and no inbound ports are opened on your network.
  • Signed probes. Each probe generates its own ed25519 key at enrolment, and GEN holds only the public half.
  • A pinned uplink. TLS 1.3, always verified, on a connection kept apart from the one that reaches your equipment.
  • Credentials held in memory only. Sealed at rest, inherited down the tree, and never written to disk by a probe.
  • Tenant isolation is enforced in the data layer, and every administrative action is audited.

Pay for What Runs

No per-sensor licence, no tiers and no prepayment. Oversight is priced per read, from £0.00001, invoiced monthly in arrears, so a simple check costs almost nothing and a heavy one costs what it takes. A probe group that is down produces no reads and costs nothing.


ExampleIndicative cost
A ping every minute, from one location£0.43 a month
An email round trip every fifteen minutes£0.17 a month
A Proxmox guest sensor, from three locations every minute£10.37 a month
A website check, from twenty locations every minute£34.56 a month

Rules, schedules, templates, escalation, state evaluation and delivery history are included and never charged. Every sensor shows its cost per read and a monthly estimate as it is configured.


See It for Yourself

Oversight has its own site at oversight.gen.uk, with an interactive demonstration of consensus voting, a cost estimator for your own estate, and a feature-by-feature comparison against the established monitoring products, drawn from each vendor's own documentation.


To talk about monitoring your estate, or moving across from an existing system, contact us and one of our engineers will take you through it.


Contact Us