ZenOps 122

ZenOps and Automotive FMEA

Automotive engineering cannot be based only on the question:

How should the vehicle work?

It must also ask:

How can the vehicle fail?

A battery can overheat.

A sensor can produce an incorrect value.

A communication network can lose messages.

A mechanical component can fracture.

Software can enter an unexpected state.

A manufacturing process can produce a defective component.

A supplier can introduce variation.

And sometimes several individually manageable failures interact to create a much larger problem.

This is why automotive engineering uses FMEA — Failure Mode and Effects Analysis.

ZenOps does not replace FMEA.

Instead, ZenOps can integrate FMEA into the complete engineering knowledge chain:

x → NDD → Requirements → ORIGIN → FMEA → StoryQ → FLEXI → Evidence → QT

Failure analysis becomes connected directly to the model of the vehicle and ultimately to the human need the vehicle exists to satisfy.


What Is FMEA?

FMEA is a structured method for asking questions such as:

  • What can fail?
  • Why can it fail?
  • What happens if it fails?
  • How serious is the consequence?
  • How likely is the failure?
  • How likely are we to detect it?
  • What controls already exist?
  • What should we do about the remaining risk?

A simplified chain is:

Object / Function
↓
Failure Mode
↓
Failure Effect
↓
Failure Cause
↓
Existing Control
↓
Risk Evaluation
↓
Action
↓
Verification

This fits naturally into ZenOps.

FMEA is fundamentally another way of exploring the relations inside the domain.


Start With the Intended Function

Before asking how something can fail, we need to understand what it is supposed to accomplish.

Suppose we have:

Object:
Battery Cooling Pump
Function:
Circulate coolant through the battery
thermal-management system.

Now we can ask:

In what ways can this function fail?

Possible failure modes include:

No coolant flow
Insufficient coolant flow
Excessive coolant flow
Intermittent operation
Incorrect rotation
Unexpected shutdown

Each failure mode represents a deviation from intended behavior.


Failure Is Relative to Need

A component failure matters because of its effect on something larger.

Suppose the cooling pump stops.

Cooling Pump Failure
↓
Reduced Coolant Flow
↓
Battery Temperature Rises
↓
Battery Power Limited
↓
Vehicle Performance Reduced

Continue upward:

Vehicle Performance Reduced
↓
Mobility Requirement Threatened
↓
NDD Need Threatened
↓
Human Need Threatened

This creates an important ZenOps principle:

A failure mode has meaning because it threatens some part of x.

FMEA should therefore not exist as an isolated spreadsheet.

It should connect back to the need structure.


Connect FMEA to ORIGIN

ZenOps ORIGIN models the vehicle as:

Objects + Relations

Consider:

Battery
cooled by
Cooling System
Cooling System
controlled by
Thermal Controller
Thermal Controller
receives data from
Temperature Sensor

Now failures can attach directly to these objects and relations.

For example:

Temperature Sensor
can fail by
Reporting Incorrect Temperature

or:

Communication Relation
can fail by
Message Loss

This is important.

Failure does not exist only inside objects.

Relations can fail too.


Relations Have Failure Modes

Suppose:

Battery Controller
communicates with
Vehicle Controller

Possible relation failures include:

  • Message not transmitted
  • Message delayed
  • Message corrupted
  • Incorrect message received
  • Communication interrupted
  • Stale information used

The controllers themselves may be functioning correctly.

The failure exists in their interaction.

Since ZenOps treats relations as first-class parts of the domain model, relation FMEA becomes natural.


Failure Modes Can Become Domain Objects

Instead of storing a failure mode only as a row in a document, ZenOps can give it identity.

For example:

FAILURE-00421
Object:
Battery Cooling Pump
Failure Mode:
Loss of coolant flow
Potential Effect:
Battery overheating
Potential Cause:
Pump motor failure

Now the failure can participate in relations.

FAILURE-00421
affects
Battery Thermal System
FAILURE-00421
threatens
REQ-00881
FAILURE-00421
detected by
Diagnostic Function
FAILURE-00421
verified by
TEST-01982

FMEA becomes part of the automotive object network.


Failure Effects Form Chains

A failure rarely stops at the component boundary.

Consider:

Temperature Sensor Failure
↓
Incorrect Temperature Value
↓
Incorrect Thermal Decision
↓
Insufficient Cooling
↓
Battery Temperature Increase
↓
Power Limitation
↓
Reduced Vehicle Performance

These are causal relations.

ZenOps can model them explicitly.

This allows engineers to ask:

What downstream effects can this failure create?

and also:

Which upstream failures could produce this observed effect?

The failure model becomes navigable in both directions.


Component Failure Versus System Effect

This distinction is critical.

A failed sensor is a component-level event.

But the customer may experience:

Vehicle power unexpectedly reduced.

The customer does not care that SENSOR-218 stopped functioning.

The customer experiences the effect.

Therefore the FMEA chain should preserve multiple levels:

Component Failure
↓
Subsystem Effect
↓
System Effect
↓
Vehicle Effect
↓
Human Effect

ZenOps keeps the technical failure connected to its consequence in reality.


FMEA Can Generate Requirements

Suppose analysis discovers:

Failure:
Cooling pump stops.
Effect:
Battery may exceed acceptable temperature.

This can generate a requirement:

The system shall detect loss of required coolant flow.

Another requirement may be:

The battery-control system shall limit power when cooling capability becomes insufficient.

Another:

A diagnostic event shall be recorded when the cooling pump fails.

Thus:

Failure Mode
↓
Required Mitigation
↓
Requirement

FMEA becomes a source of requirements.


Requirements Can Generate FMEA Questions

The relationship also works in the opposite direction.

Suppose we have:

The vehicle shall maintain controllability during braking.

FMEA can ask:

What failures could prevent this requirement from being satisfied?

Possible answers include:

  • Wheel-speed sensor failure
  • Brake actuator failure
  • Controller failure
  • Communication failure
  • Power-supply failure
  • Incorrect software state

Therefore:

Requirement
↓
What Can Prevent This?
↓
Failure Modes

Requirements and FMEA reinforce each other.


Failure Detection Is Not Failure Prevention

This distinction matters.

Suppose the system detects a failed sensor.

That does not necessarily prevent the failure.

Instead, detection allows the system to respond.

The full chain may be:

Failure
↓
Detection
↓
Isolation
↓
Degraded Operation
↓
Driver Notification
↓
Recovery / Service

Different requirements may be needed for each stage.


FMEA and Automotive Patterns

Many failure-response structures repeat.

A useful pattern might be:

Detect → Isolate → Degrade → Report → Recover

For example:

Sensor Failure
↓
Detect Invalid Signal
↓
Ignore Failed Sensor
↓
Use Degraded Control Strategy
↓
Record Diagnostic Event
↓
Recover or Request Service

This can become part of the ZenOps automotive Pattern Library.

Future systems can reuse the pattern.


Pattern Libraries Can Contain Failure Knowledge

An engineering pattern should not contain only:

Here is how to build this.

It can also contain:

Here is how this typically fails.

For example:

Sensor Pattern
│
├── Intended Function
├── Architecture
├── Interfaces
├── Known Failure Modes
├── Detection Patterns
├── Degraded Modes
├── StoryQ Scenarios
└── Verification Methods

Now decades of failure knowledge can accumulate around reusable engineering structures.


FMEA Generates StoryQ/Gherkin Scenarios

This is where the previous parts of the ZenOps chain connect strongly.

Suppose FMEA identifies:

Wheel-speed sensor signal unavailable.

That failure mode can become a StoryQ scenario:

Scenario: Wheel-speed sensor signal becomes unavailable
Given the vehicle is moving
And all wheel-speed signals are initially valid
When one wheel-speed signal becomes unavailable
Then the system shall detect the failed signal
And the affected control function shall enter the defined degraded mode
And the diagnostic event shall be recorded
And no unsafe control output shall be generated

The FMEA entry has become executable behavior.


Every Important Failure Mode Should Ask for Evidence

A failure analysis is incomplete if it merely says:

We have mitigation.

ZenOps asks:

What evidence demonstrates that the mitigation actually works?

The chain becomes:

Failure Mode
↓
Mitigation Requirement
↓
StoryQ Scenario
↓
Failure Injection
↓
Observed Response
↓
Evidence

This transforms FMEA from prediction into verification.


Failure Injection Makes FMEA Real

Suppose the analysis says:

If the coolant pump fails, the system detects the failure and limits battery power.

Test it.

Physically disconnect the pump.

Simulate the electrical failure.

Inject the diagnostic condition.

Interrupt communication.

Then observe the system.

Did it detect the failure?

How quickly?

Did power limitation occur?

Was the correct diagnostic event recorded?

Did the vehicle remain safe?

Now the FMEA has met reality.


FMEA Generates FLEXI Micro-Sprints

Each important unresolved failure mode can become a FLEXI work item.

For example:

FLEXI-0921
Question:
Does the thermal-control system correctly
respond to loss of coolant flow?
Given:
Representative operating condition
Action:
Inject pump failure
Expected:
Failure detected
Battery protected
Diagnostic recorded
Output:
Evidence

One FMEA row can therefore generate one or more targeted micro-sprints.


Prioritize Uncertainty, Not Paperwork

Traditional FMEA exercises can become large tables containing hundreds or thousands of entries.

ZenOps should resist turning this into documentation for its own sake.

The valuable question is:

Which failure modes currently represent the greatest unresolved uncertainty or unacceptable risk?

Those should generate work first.

High Risk + Low Evidence
↓
FLEXI
↓
Test
↓
Evidence
↓
Updated Risk

FMEA becomes an execution driver.


FMEA and Quality Thresholds

Failure analysis should contribute directly to QT.

For example:

BATTERY MODULE QT
[ ] Critical failure modes identified
[ ] Effects evaluated
[ ] Causes investigated
[ ] Detection mechanisms implemented
[ ] Mitigations implemented
[ ] Critical failure scenarios tested
[ ] Residual risks evaluated
[ ] Evidence accepted

The module should not cross QT merely because the FMEA document exists.

It crosses when sufficient evidence supports the risk controls.


FMEA Status Can Become Evidence-Based

Instead of:

FMEA complete: YES

ZenOps can expose:

Cooling Pump Failure
Identified: PASS
Detection:
PASS
Degraded Operation:
PASS
Diagnostic Reporting:
PASS
Recovery:
PARTIAL
Physical Validation:
UNKNOWN

This tells the project what remains uncertain.


UNKNOWN Is Valuable

Suppose:

Physical Validation: UNKNOWN

That is not administrative failure.

It is useful information.

UNKNOWN generates a question.

The question generates FLEXI work.

The work generates evidence.

UNKNOWN
↓
Question
↓
Experiment
↓
Evidence
↓
Updated FMEA

This is the ZenOps learning loop again.


Severity Matters

Not every failure deserves equal effort.

A broken cupholder and a braking-control failure do not have the same consequences.

FMEA therefore considers the severity of the effect.

ZenOps can connect severity to the human impact.

For example:

Failure
↓
Vehicle Effect
↓
Human Consequence
↓
Severity

This prevents technical scoring from becoming disconnected from reality.


Occurrence Matters

Another question is:

How likely is this failure?

Evidence may come from:

  • Component reliability data
  • Supplier history
  • Testing
  • Simulation
  • Previous vehicle programs
  • Field data

Occurrence should therefore evolve as evidence accumulates.

A failure considered rare during concept development may prove more common in fleet operation.

The model should be allowed to change.


Detection Matters

FMEA also asks whether a failure is likely to be detected before it causes harm.

For example:

Sensor Failure
↓
Diagnostic Detection
↓
Driver Warning
↓
Service Action

If the failure is difficult to detect, risk may remain higher.

This can generate new requirements for diagnostics or monitoring.


Risk Priority Is a Decision Aid, Not Reality

FMEA methods often use structured ratings or action priorities to help decide where engineering attention is needed.

ZenOps can use such prioritization.

But the number should not replace engineering reasoning.

Two failure modes with similar scores may have very different consequences, uncertainty, or evidence quality.

The important questions remain:

What can happen?

Why?

How serious is it?

What evidence do we have?

What should we do next?


Design FMEA and Process FMEA

Automotive FMEA applies not only to the vehicle design.

It also applies to manufacturing.

We can distinguish broadly between:

Design FMEA — DFMEA

and:

Process FMEA — PFMEA

DFMEA asks:

How can the product design fail?

PFMEA asks:

How can the manufacturing process fail to produce the intended product?

ZenOps can integrate both into the same domain.


Example DFMEA

Consider a battery connector.

Possible design failure:

Object:
High-Voltage Connector
Failure Mode:
Electrical contact lost
Possible Effect:
Loss of propulsion power
Possible Causes:
Mechanical separation
Contact degradation
Thermal damage

This can generate requirements for:

  • Mechanical retention
  • Contact monitoring
  • Fault detection
  • Safe power-down behavior

And each requirement can generate evidence.


Example PFMEA

Now consider installation of the same connector.

Manufacturing Operation:
Install High-Voltage Connector
Failure Mode:
Connector not fully seated
Effect:
Intermittent electrical contact
Cause:
Incorrect assembly
Misalignment
Insufficient verification

Potential controls may include:

  • Mechanical poka-yoke
  • Position detection
  • Electrical verification
  • End-of-line testing

Now manufacturing FMEA connects directly to the same physical object.


DFMEA and PFMEA Should Connect

This is a major opportunity for a domain-model approach.

Suppose DFMEA identifies:

Partial connector engagement can create dangerous behavior.

PFMEA should know that this failure is important.

The manufacturing process should therefore include controls specifically designed to prevent or detect it.

DFMEA Failure
↓
Critical Product Characteristic
↓
PFMEA
↓
Manufacturing Control
↓
Inspection
↓
Production Evidence

The engineering and factory risk models become connected.


FMEA Can Generate Manufacturing Tests

Suppose PFMEA identifies:

Fastener may receive insufficient torque.

That can generate:

Requirement:
Torque must remain within specified limits.
Manufacturing Control:
Controlled torque tool.
Evidence:
Recorded torque result.

For a physical vehicle:

Vehicle #000142
↓
Fastener #F-0081
↓
Torque Measurement
↓
PASS

Risk analysis now connects all the way to vehicle-instance evidence.


Supplier FMEA Can Join the Same Network

Suppliers may perform their own design and process FMEA.

Rather than receiving only a PDF, ZenOps can conceptually connect relevant supplier risks to:

  • Supplied components
  • Requirements
  • Interfaces
  • Manufacturing controls
  • Incoming inspection
  • Vehicle tests

The supply chain becomes part of the same risk model.


Software Failure Modes Matter Too

Modern vehicles contain enormous amounts of software.

Software-related failure modes can include:

  • Incorrect state transition
  • Timing failure
  • Incorrect calculation
  • Missing message
  • Stale data
  • Memory/resource exhaustion
  • Recovery failure
  • Configuration mismatch

These failures can participate in the same ZenOps structure:

Software Object
↓
Failure Mode
↓
System Effect
↓
Requirement
↓
Scenario
↓
Test
↓
Evidence

Hardware and software risk become part of one system model.


FMEA Should Follow Interfaces

Suppose:

Sensor
sends
Measurement
to
Controller

Ask failure questions about the relation:

What if the message never arrives?
What if it arrives late?
What if it contains an invalid value?
What if it freezes at the previous value?
What if the controller interprets it incorrectly?

This systematically explores interaction risk.

Because ZenOps models relations explicitly, interface FMEA can become especially powerful.


Failure Chains Can Reveal Common Causes

Suppose several systems depend on one power source.

Power Supply
│
├── Sensor A
├── Sensor B
└── Controller C

Individual FMEAs may treat the systems separately.

But the object network reveals a shared dependency.

Failure of the power supply could disable all three simultaneously.

This exposes a common-cause risk.

Network thinking can therefore complement traditional component-by-component analysis.


The Domain Model Can Help Discover FMEA Scope

If the complete automotive domain model contains:

  • Objects
  • Relations
  • Functions
  • Interfaces
  • Requirements

then FMEA scope can be generated systematically.

For every important object:

How can this object fail?

For every important relation:

How can this interaction fail?

For every important function:

How can this function be absent, incorrect, excessive, delayed, or unintended?

For every important requirement:

What could prevent this requirement from being satisfied?

This gives FMEA a structural foundation.


Patterns Can Suggest Failure Modes Automatically

Suppose the Pattern Library knows the pattern:

Sensor
↓
Communication
↓
Controller
↓
Actuator

The pattern library may also know common failure categories:

Sensor:
No signal
Incorrect signal
Noisy signal
Frozen signal
Communication:
Missing message
Delayed message
Corrupted message
Controller:
Incorrect decision
No decision
Late decision
Actuator:
No actuation
Partial actuation
Unexpected actuation

When the pattern is instantiated, candidate FMEA entries can be suggested.

Engineering experience becomes reusable.


FMEA Becomes Organizational Memory

Every vehicle program discovers new failure modes.

Without structured reuse, the next program can repeat old mistakes.

ZenOps can preserve:

Failure Pattern
↓
Known Causes
↓
Known Effects
↓
Successful Controls
↓
Failed Controls
↓
Verification Scenarios
↓
Field Evidence

The FMEA becomes more than a project deliverable.

It becomes accumulated engineering knowledge.


Field Failures Must Feed Back Into FMEA

Development teams cannot predict every failure.

Reality will discover some for us.

Suppose field vehicles reveal a failure mode not present in the original analysis.

The loop should be:

Field Failure
↓
Diagnostic Investigation
↓
Root Cause
↓
New / Updated FMEA
↓
Requirement Update
↓
StoryQ Scenario
↓
Corrective Work
↓
Verification
↓
Regression Evidence

The FMEA remains alive.


Every Serious Field Failure Should Leave Knowledge Behind

A repaired customer vehicle is not enough.

The organization should ask:

What have we learned that prevents this failure from surprising us again?

The answer may include:

  • New FMEA entry
  • New pattern
  • New requirement
  • New diagnostic
  • New test
  • New manufacturing control
  • New supplier requirement

The failure becomes organizational memory.


FMEA Changes as Evidence Changes

Suppose a failure was originally believed to be extremely rare.

After 100,000 vehicles enter service, field evidence shows otherwise.

The occurrence assessment changes.

Likewise, a detection mechanism believed to be highly effective may prove less effective in reality.

FMEA should therefore not be frozen at production launch.

It should evolve with evidence.


The Fleet Becomes an FMEA Laboratory

Once vehicles operate in the field, the organization gains enormous amounts of real-world information.

Potential patterns may appear:

Failure X
occurs mainly in
Climate Y
Failure A
correlates with
Software Version B
Failure C
correlates with
Supplier Batch D

These relations can update risk models.

The fleet becomes part of the evidence system.


FMEA and the Quality Threshold Loop

We can now combine the pieces:

Domain Object
↓
Function
↓
Failure Mode
↓
Effect
↓
Cause
↓
Risk
↓
Mitigation Requirement
↓
StoryQ Scenario
↓
FLEXI Work
↓
Test
↓
Evidence
↓
QT

If evidence is insufficient:

QT
├── PASS → Accept current risk
├── PARTIAL → More evidence
├── FAIL → Redesign / Correct
└── UNKNOWN → Investigate

The loop continues.


From FMEA Spreadsheet to Failure Knowledge Network

Traditional FMEA is often represented as rows and columns.

That representation remains useful.

But ZenOps adds another view.

Imagine:

Cooling Pump
│
├── can fail as → No Flow
│ │
│ ├── causes → Battery Heating
│ │
│ ├── detected by → Flow Diagnostic
│ │
│ ├── mitigated by → Power Limitation
│ │
│ └── verified by → TEST-812
│
└── manufactured by → Supplier A

Now the failure is connected to the rest of the engineering model.

FMEA becomes a network rather than an isolated artifact.


Trace Failure All the Way to x

The ultimate traceability chain might become:

Human Need
↓
NDD
↓
Requirement
↓
System Function
↓
Object / Relation
↓
Failure Mode
↓
Failure Effect
↓
Mitigation
↓
StoryQ Scenario
↓
Test
↓
Evidence
↓
QT

This tells us both:

why the system exists

and:

what happens when it stops doing what it exists to do.


FMEA Is the Negative Image of the Domain Model

There is an interesting way to think about this.

The normal domain model asks:

How should the vehicle work?

FMEA asks:

How can that intended structure break?

If the domain model says:

Sensor
reports to
Controller

FMEA asks:

What if it does not?

If the requirement says:

Battery temperature shall remain within range.

FMEA asks:

What could make it leave that range?

If the architecture says:

This module provides braking control.

FMEA asks:

What happens when it cannot?

FMEA is therefore almost a negative image of the intended system.

Together, the two models provide a more complete understanding.


Design for Failure, Not Only Success

A vehicle that works only when everything works perfectly is not a robust vehicle.

Real systems must expect:

  • Components to degrade
  • Sensors to fail
  • Messages to disappear
  • Humans to make mistakes
  • Manufacturing variation to occur
  • Environmental conditions to become extreme

Good automotive engineering therefore does not merely design the success path.

It designs the failure paths.

ZenOps can make those paths explicit.


From Failure Prediction to Evidence

The most important transformation is this:

We think this could fail.
↓
We understand the effect.
↓
We design a response.
↓
We implement the response.
↓
We deliberately create the failure.
↓
We observe what happens.
↓
We collect evidence.

Now FMEA has moved beyond analysis.

It has become part of the engineering execution system.


The Complete ZenOps FMEA Loop

The full loop can be summarized as:

x
↓
NDD
↓
Requirements
↓
ORIGIN
↓
Objects + Relations
↓
Functions
↓
FMEA
↓
Failure Modes
↓
Effects + Causes
↓
Risk
↓
Mitigation
↓
StoryQ / Gherkin
↓
FLEXI
↓
Failure Injection
↓
Evidence
↓
QT
↓
Vehicle
↓
Field Evidence
↓
Updated FMEA

The process never truly ends.

Reality keeps teaching the model.


Failure Is Information

The purpose of FMEA is not to imagine that every possible failure can be eliminated.

That is impossible.

The purpose is to understand failure sufficiently well to make better engineering decisions.

ZenOps extends that idea.

A failure mode becomes a question.

The question generates a requirement.

The requirement generates a scenario.

The scenario generates a test.

The test generates evidence.

The evidence informs a Quality Threshold.

And field experience eventually challenges everything we believed.

This creates a different relationship with failure.

Failure is not merely something to avoid.

It is something to model, test, observe, learn from, and remember.

The automobile becomes safer and more robust not because engineers assume that everything will work.

It becomes safer because engineers systematically ask:

What happens when it doesn’t?

That is the connection between ZenOps and automotive FMEA.

Model success. Model failure. Test both. Preserve the evidence. Learn from reality.

Leave a comment