Findings: what broke, and whether it can wait
Incident triage is deciding, in the first minutes, what broke, how bad it is and whether it can wait; on DigitalOcean the answer is spread across droplet graphs, App Platform activity and the status page. Duskwatch is an independent iPhone app for DigitalOcean, in development; it gathers that answer into findings, puts what needs attention first and says why.
Updated
A finding is more than a status
A finding is one plain sentence about one resource that says what is wrong, since when, and what it is based on. "Active" tells you a droplet is running. "Disk full in ~5 days at this rate" tells you what to do this week. Findings are worked out on your iPhone from what the app reads from DigitalOcean for the screen you are on.
Latest deploy failed · api-prod
Likely cause: api didn't pass its health check
Still live: 3f2a91c · since 22:40
The likely cause is labelled as likely. It comes from DigitalOcean's own reason for the failed step when there is one, from log lines you have already opened, or from the stage that failed. DigitalOcean's raw reason is one tap away, so you can check the guess.
Findings by resource
| Resource | Finding | What it tells you |
|---|---|---|
| Droplet (web-01) | SSH is open to the internet | Its cloud firewall lets port 22 in from any address |
| Droplet (web-01) | Disk full in ~5 days at this rate | The disk trend reaches 100% within days |
| App (api-prod) | worker keeps restarting | A component is in a crash loop |
| App (api-prod) | Latest deploy failed | With a likely cause, and the version still live |
| Kubernetes (k8s-prod) | 1 node not ready for 25 min | A node in the cluster cannot take work |
| Load balancer (lb-web) | 1 of 2 droplets behind it is off | Half the capacity behind it is gone |
| Certificate (static-example) | Certificate expires in 9 d | Renew it, or check that renewal works |
| Volume (scratch-01) | Not attached, still billed | Storage you pay for and do not use |
| Spaces | 1 full-access key can reach every bucket | One leaked key exposes all buckets |
Findings on Kubernetes, load balancers, volumes, Spaces and certificates stay in the app. Only the events listed on the push alerts page wake you.
Needs attention first
The overview sorts by consequence, not by resource type: something down comes before something filling up, which comes before housekeeping. Routine events stay quiet, so a single reboot this week does not compete with a failed deploy. With more than one team, the team switcher shows a problem count beside any team that needs you, so you can skip the rest.
The disk estimate is deliberately cautious. It only looks at a disk that is already more than 70% full, reads its last week of usage, fits a trend that ignores the saw-tooth of log rotation, and says nothing when the trend is too noisy to trust. Charts show missing data as gaps instead of drawing a line through them.
Droplet graphs for CPU, load average and memory come from DigitalOcean's metrics agent, which runs on the droplet.
DigitalOcean docs: track droplet performance (external link), verified
App Platform crash logs record the instance's output before it crashed and a potential reason for the crash.
DigitalOcean docs: App Platform logs (external link), verified
App Platform can alert on a failed deployment and on restart count, by email or Slack.
DigitalOcean docs: App Platform alerts (external link), verified
Is it DigitalOcean, or is it me?
Half of a night's triage is ruling out the platform. The app reads DigitalOcean's status and matches incidents and maintenance to the regions and products your team uses. An incident in a region you do not use stays out of the way; one that touches web-01's region sits next to the finding, so you can see both at once.
DigitalOcean's status page lists components by product and by data center region, such as FRA1 or NYC3.
DigitalOcean status page (external link), verified
More on this on the platform status page.
From finding to action, or to a hand-off
A finding ends in a next step. Some are a guarded action on the phone: reboot web-01, roll back api-prod. Others belong somewhere else, and the app gets you there quickly.
- Share a log excerpt with a teammate, with the app, component and time range in its header.
- Open the droplet in your SSH app, or copy its IP.
- Open the resource in DigitalOcean's control panel to change what the app does not.
- Take a guarded action: one sheet, a swipe for anything disruptive, then Face ID.
What Duskwatch doesn't do here
- No root-cause analysis: a likely cause is a labelled guess, with DigitalOcean's own reason beside it.
- No pods, containers or processes: findings stop at the droplet, app, cluster and node level.
- No custom finding rules or thresholds of your own.
- No push for findings on Kubernetes, load balancers, volumes, Spaces or certificates.
- No CPU, memory or load charts for a droplet without DigitalOcean's metrics agent.
Questions
What is a finding?
A finding is one plain sentence about one resource that says what is wrong, since when, and what it is based on, such as "Disk full in ~5 days at this rate". Findings are worked out on your iPhone from data the app already reads.
How does it tell a disk is filling?
For a disk more than 70% full, it reads the last week of usage, fits a trend that ignores log-rotation dips and projects when it reaches 100%. When the trend is too noisy, it says nothing.
Is it DigitalOcean, or is it my app?
Duskwatch matches DigitalOcean's status to the regions and products you use and shows a relevant incident next to the finding. If nothing matches, the problem is most likely on your side.
Is the likely cause of a failed deploy always right?
No, which is why it is labelled as likely. It is based on DigitalOcean's reason for the failed step, log lines you have opened or the stage that failed, and the raw reason is one tap away.
Can I share what I found?
Yes. You can share a log excerpt, open the droplet in your SSH app or open the resource in DigitalOcean's control panel. Nothing leaves the app until you tap Share.