←All Products APPLICATIONS · AI APPS

Nightwatch

AI operations engineer

No more manual toil: automatic inspection and diagnosis

Nightwatch is your digital operations engineer. AI agents take over routine monitoring, inspections, releases and troubleshooting, unattended around the clock. You make the key decisions; Nightwatch does the rest.

90%+Repetitive work automatedInspections, releases and troubleshooting
30minMean time to recoveryDown from hours to minutes
80%+Fewer human errorsTiered approvals and full audit trail
24×7Always on dutyNight-time incidents heal themselves
PAIN POINTS

Six pain points, taken apart

Alert storms, night-time firefighting, repetitive inspections, troubleshooting by memory, risky releases and endless reporting — the six core pains of traditional operations, and how Nightwatch solves them.

70%+ false alarms

Alert storms

The alerting platform floods all day and seven in ten alerts are false. You get up in the middle of the night to check each one, only to find nothing — while real incidents drown in the noise and take far longer to fix.

3 all-nighters a month

Firefighting at night

A call at 3 a.m., a night spent digging through logs and metrics, the root cause found at dawn — and a normal working day ahead. Always on standby, judgment and efficiency suffer, and haste breeds mistakes.

2 hours of toil a day

Repetitive inspections

Two hours a day spent staring at dashboards, logs and disks. And manual inspections can only happen so often — problems break out between rounds.

Knowledge only in people's heads

Knowledge walks out the door

Senior engineers' troubleshooting know-how lives in their heads and leaves when they do. Newcomers start from scratch and fall into the same holes; team knowledge never builds up.

A late-night release every 2 weeks

Releases on a tightrope

Every release is nerve-racking: back up, roll out gradually, watch metrics for hours, roll back overnight if something breaks. All the risk sits with the person on call, and one wrong move can cause a major outage.

3 hours of reports a week

Reports that never end

Weekly and monthly reports, post-mortems and security reports written late into the night, with data compiled by hand and definitions that never quite match.

OPS LIFECYCLE

From alert to resolution, fully automated

From the asset inventory to the post-mortem, across the whole operations lifecycle. The system moves each step forward; you decide only at the key points.

  1. Asset inventory

    All infrastructure under management

  2. Smart inspection

    Finding risks 24/7

  3. Alert

    Diagnosis within 30 seconds

  4. Diagnosis

    Root cause from correlated signals

  5. Tiered self-healing

    Safe actions close the loop

  6. Change & release

    Gradual rollout, automatic rollback

  7. Post-mortem

    Written and sent automatically

CORE MODULES

Twelve modules across the operations lifecycle

From monitoring and inspection to vulnerability management and reporting, AI is involved at every step.

Inspection & monitoring

Replaces the manual morning check, automatically inspecting resources, applications, logs, middleware and business flows to catch problems early.

ResourcesApplicationsLog analysisAlert noise reduction

Incident response & self-healing

Diagnosis starts within 30 seconds of an alert, correlating signals to find the root cause; safe actions run automatically, risky ones go to a person for approval.

Auto-diagnosisTiered self-healingIncident briefsKnowledge capture

Change & release

Automatic pre-release checks, gradual rollout and post-release verification, with automatic rollback if anything goes wrong.

Canary releasePre-release checksPost-release checksOne-click rollback

Natural-language ops

"Restart Nginx on the web cluster." "Show the top 10 disks by usage." One sentence and it's done — no commands to remember.

One-sentence opsBatch operationsDatabase opsConfig management

Asset management

Servers, containers, services and network resources under one roof, with automatic discovery and reconciliation so records match reality.

ServersContainer clustersAuto-discoveryChange tracking

Capacity & performance

Forecasts resource levels from historical trends and warns of overload early — from firefighting to prevention.

Capacity forecastingResource optimizationPerformance analysisSlow-query finding

Security operations

Automatic triage of security alerts, key expiry reminders and rotation, and regular baseline checks, so attacks and risks are visible and can be dealt with.

Alert triageKey auditBaseline checksAccess audit

Vulnerability management

Vulnerabilities go from scan results to a closed loop: found, assessed, remediation planned, approved and executed, then retested.

Auto-ingestRisk ratingRemediation plansRetesting

Reporting center

Inspection reports, post-mortems and weekly or monthly reports in one click, with chart-based versions for leadership and customers.

Inspection reportsPost-mortemsWeekly & monthlyExternal versions

Knowledge · system brain

Runbooks stored in structured form, troubleshooting experience captured as a team asset and retrieved automatically during incidents — it doesn't leave when people do.

Runbook libraryArchitecture Q&ASystem Q&ALong-term memory

Collaboration & integration

Conversational ops in WeCom, DingTalk, Feishu and Slack, with incident alerts and approvals wherever you are — operate where you work.

Chat opsIncident alertsApprovalsWeb console

Platform & compliance

RBAC permissions, platform configuration, every operation logged for 180 days and complete, exportable approval chains to meet compliance requirements.

RBACAudit trailApproval chainData export
DESIGN PRINCIPLES

Safety before speed; every action auditable

Four core principles run through every module: read-only by default to minimize production impact; safe actions close the loop automatically; risky actions require human confirmation (human-in-the-loop); and everything is logged, traceable and reversible. No feature trades safety for speed.

01Read-only whenever possible
02Self-heal whenever safe
03People confirm risky actions
04Every action is auditable
USE CASES

Real operations work, handed to Nightwatch

From unattended inspections to night-time self-healing, from one-sentence ops to instant reports.

Daily 08:00 · automatic

Unattended inspection

Inspects every core system and sends the report before 9:00, with anomalies flagged in red so risks are caught early.

03:12 · within 30 seconds

Night-time self-healing

An alert triggers an agent: automatic diagnosis → safe self-healing → a note to the on-call group, with no one involved.

Web / IM / voice · instant

One-sentence ops

"Restart Nginx on the web cluster." "Show the top 10 disks by usage." Executed straight away, with results shown in real time.

Release window · end to end

Safe releases

Pre-release checks → gradual rollout → post-release verification, with a 15–30 minute watch period and automatic rollback.

Critical · 72 hours

Closing out vulnerabilities

A critical vulnerability found in a scan is assessed for impact automatically, a fix is planned, executed after approval and retested — escalated automatically if it runs late.

Month end · one click

Instant reports

Weekly and monthly reports, post-mortems and security reports in one click, with charts, consistent definitions and automatic delivery.

MODEL ROUTING

Tiered model routing, controlled cost

Routine tasks use low-cost models; hard incidents switch to stronger reasoning models. Multiple models race and back each other up for low latency and high availability. Domestic models by default, with full on-premises deployment.

Main

DeepSeek

Day-to-day inspection, diagnosis and reporting.

Main

Qwen

Backs up DeepSeek, racing and failing over by task.

Hard incidents

Stronger reasoning models

Complex incidents switch to a stronger reasoning model automatically.

Pluggable

Your own models

The model layer is pluggable, including models you deploy privately.

More Applications

MEDIA & CONTENT

tide CMS

Content publishing system

Manage content for websites, apps, IPTV and OTT from one platform — edit once, publish to every screen.

MEDIA & CONTENT

tide VMS

Media asset management

One repository for video, audio, text, images and specialist data, with transcoding, editing, watermarking, subtitles and multi-channel output.

LIVE & VIDEO

LiveSphere Dream

Live streaming management

Visually schedule TV, online and ad-hoc live channels, with multiple stream formats in and out for web, IPTV and mobile.

MEDIA & CONTENT

tide APP

Media app builder

Build media apps with low-code and ship to multiple platforms quickly, with live, on-demand, news, UGC and interactive features.

LIVE & VIDEO

LiveSphere Air

Real-time recording and clipping

Records live signals continuously so clips can be cut and published about a minute after airing, with frame-accurate editing in the cloud.

LIVE & VIDEO

Juxian Cloud

Media cloud service

Cloud live streaming, on-demand, paid content and interaction for major events and corporate training, with no infrastructure to run.

APP BUILDING

tide JuDA

Low-code app builder

Drag-and-drop interfaces and workflows so managers, developers and business staff can build business apps quickly, with data shared across systems.

USERS & DATA

tide UIC

User interaction center

Unified sign-in and user management, interactions pooled from WeChat and Weibo, plus voting, comments, favorites and cross-screen features.

USERS & DATA

tide DAM

User analytics

Tracks user behavior across websites, video, apps and WeChat, supporting user profiles, personalized recommendations and targeted outreach.

USERS & DATA

tide Cloud RC

Cloud content collection

Collects third-party articles, video and specialist data by rule, maps it into your repository, and monitors every collection job.

PLATFORM ADMIN

tide PMC

Platform management center

One entry point for users and permissions across business systems, with full logging, health monitoring, security and unified upgrades.

Want to see it in your own business?

Get Started All Products