Model and scenario reference¶
system.model.yaml¶
meta:
name: storefront # required
version: "1.0" # required, quoted
providers: [mgt-tool/kubernetes@>=3.0.0]
vars: { namespace: production } # substituted as {namespace} in probes
strict_types: false # true: an untyped component is an error
scenarios: none # opt out of scenarios.yaml and its drift check
components:
api:
type: kubernetes.deployment # required: <provider>.<type>, or bare <type>
resource: shop-api-{env} # real resource name; {var} placeholders allowed
source: discovered # written by `model build`; absent = authored, never deleted by build
providers: [mgt-tool/aws] # override meta.providers for this component
vars: { namespace: payments } # override meta.vars
depends:
- on: rds
- on: vault
while: vault.state == "starting" # edge exists only while true
- on: [web-a, web-b]
need: 1 # redundancy group: holds while ≥1 is healthy
healthy:
add: [restart_count < 3] # the type's rules AND these
failure_modes: # extra propagation the type doesn't declare
degraded: { can_cause: [upstream_5xx] }
A probe reads resource: when it is set, else the component's key, and a command's {name} is what it reads. A Deployment and the Service in front of it often share a name: key them by kind, deployment/shop-api and service/shop-api, and probes read shop-api for both, as kubectl get service/shop-api does. The prefix must be the component's type; svc/shop-api reads whole, and validate says so. Expressions name such a component whole: service/shop-api.endpoint_count > 0.
source: says where a component came from. mgtt model build writes source: discovered on what it found. Any other component is authored: business processes, external services, hand-written wiring. A rebuild keeps authored components whether discovery returns them or not, and names them as kept. Only a discovered component that discovery stops returning goes through the deletion gate (--allow-deletes, --tombstone).
Health rules¶
healthy: has three forms:
add: […]keeps the type's default rules and also requires yours.replace: […]uses only your rules.- A bare list also replaces the type's rules, and
validatewarns about each type rule it drops. A redis rule ofcache_hit_ratio > 70alone would no longer checkavailable == true.
Rules have the form <fact> <op> <value>, with ==, !=, <, >, <= or >=. All rules must hold. Values can be numbers, booleans or quoted strings. A bare word is read as follows:
- against a string fact, it's a literal (
phase == Bound); - otherwise, if it's a fact of the same component, it means that fact (
ready_replicas == desired_replicas); - otherwise it's a variable, looked up in the component's
vars, thenmeta.vars, then the provider's default.
A variable that's set nowhere leaves the rule undecided, and validate warns.
A type says what healthy means twice: in its rules, and in which of its states is the default. If an override makes them disagree, validate warns and gives the facts where it happens. Rules that are loosened so they still hold in saturated mean simulate (which reads states) and diagnose (which reads rules) disagree there. When the divergence is deliberate, list those states:
opensearch:
type: deployment
healthy:
replace: [ready_replicas >= 1]
healthy_diverges_from: [crashed, degraded] # one ready replica serves search on stage
Disagreements in a type's own defaults are the provider's to fix; mgtt provider validate reports them.
Dependencies¶
api depends on rds means a broken rds can break api. Use while: for a conditional edge, such as a blue/green service that follows its live color (while: selector_value == blue). A redundancy group with need: holds while at least that many members are proven healthy. A member that couldn't be read doesn't count. Active/passive pairs are not groups: model them with while:.
Several models¶
One repo can hold several models, for example edge.model.yaml and data.model.yaml. Each command takes --model <path>. Models do not import one another.
Scenarios: scenarios/*.yaml¶
name: rds forbidden
description: IAM denies every rds probe; api is crashing.
entry: nginx # where diagnosis starts; default: the model's entry point
inject: # component → fact → value
api: { ready_replicas: 0, desired_replicas: 3, restart_count: 9 }
unresolved: # probes that ran but produced no value
rds: { available: forbidden, connection_count: forbidden }
expect:
root_cause: api # required; `none` for all-healthy
path: [nginx, api] # ordered subsequence of the actual path
eliminated: [frontend] # subset of the actual list
cannot_rule_out: [rds] # left undecided by forbidden or transient facts
not_eliminated: [rds] # must stay in play
redundancy_degraded: [web-a] # broken members their group absorbed
unresolvedoutcomes areforbidden,transient(both unknown, so the component is kept in play) ornot_found. If every fact of a component isnot_found, the component is absent, and it can be the root cause.- Inject enough facts for the state you mean.
ready_replicas: 0withoutrestart_countcan resolve todegradedrather thancrashed. - A component you don't mention has no facts, so it is never eliminated.
entrynames the component diagnosis starts from, asplan --componentdoes. Set it for a failure the model's entry point can't reach, such as a background job no request path touches.
Run them with mgtt simulate --all, or --scenario <file> for one. --from-scenarios checks every enumerated chain. --fuzz N checks that the engine reaches a conclusion from random, partial evidence.
scenarios.yaml (generated)¶
mgtt model validate --write-scenarios writes every failure chain the model allows, stored as a failure graph (format: graph/v1). Don't edit it. Regenerate it and commit it. validate fails once the model would produce a different graph; comments and version bumps don't count. validate --check-scenarios runs only that check.