Tag: Solution Assurance

  • The Tao of Solution Assurance

    Solution Assurance is the ancient corporate discipline of examining a design after everybody important has already committed to it, identifying several serious risks, documenting them carefully, and then watching the programme proceed exactly as before.

    It is sometimes confused with governance.

    This is unfair to governance.

    Governance occasionally stops things.

    Solution Assurance exists in the delicate philosophical territory between architecture, risk management, quality control and ritual sacrifice.

    Its purpose is simple:

    To provide confidence.

    Not necessarily correctness.

    Not necessarily safety.

    Certainly not certainty.

    Confidence.

    Confidence is extremely important because executives become nervous when presented with reality.

    A red RAG status creates anxiety.

    A green RAG status creates confidence.

    An amber RAG status creates meetings.

    Therefore the experienced Solution Assurance practitioner understands the first teaching:

    The colour is not the risk.

    The colour is what management can emotionally tolerate this week.

    1. Assurance Begins After the Decision

    In theory, assurance should begin early.

    In practice, the sequence is:

    1. Vendor selected.
    2. Contract negotiated.
    3. Budget announced.
    4. Programme mobilised.
    5. Launch date communicated.
    6. Solution designed.
    7. Assurance invited.

    This is efficient because it prevents inconvenient technical facts from interfering with commercial momentum.

    You will receive an email.

    Subject: Architecture Assurance – Urgent

    It will say:

    Hi,

    Could you please provide assurance against the attached design? We are seeking approval at Thursday’s board.

    This should be straightforward as the solution has already been reviewed extensively by the supplier.

    Thanks.

    It is Wednesday afternoon.

    The attachment is 184 pages.

    Half the diagrams are unreadable.

    The security section says:

    Security requirements will be confirmed during implementation.

    The implementation began three months ago.

    Welcome to Assurance.

    2. “Assured” Does Not Mean “Good”

    A solution can be:

    technically questionable,

    operationally immature,

    poorly documented,

    expensive,

    dependent on unsupported technology,

    and nevertheless assured.

    This is because assurance is rarely a binary judgement.

    It is more sophisticated.

    You can say:

    Assured subject to conditions.

    This is one of the great phrases of corporate civilisation.

    It means:

    We have identified several reasons why this may go badly wrong, but everyone has a steering committee in ten minutes.

    The conditions are then recorded.

    The programme acknowledges them.

    The programme continues.

    Six months later, during the incident review, someone asks:

    “Why wasn’t this risk identified?”

    You produce the assurance report.

    Page 17.

    Risk 4.

    Highlighted.

    Bold.

    Rated Red.

    The room becomes quiet.

    This is the closest Solution Assurance gets to physical pleasure.

    3. Never Ask “Is It Secure?”

    This is not a useful question.

    Everything is secure in PowerPoint.

    Ask:

    How is authentication implemented?

    How are privileged accounts controlled?

    Where are secrets stored?

    What is logged?

    Who can access the logs?

    How is encryption implemented?

    Who owns the keys?

    How are certificates rotated?

    What happens if identity is unavailable?

    How are vulnerabilities patched?

    What is exposed externally?

    How is compromise detected?

    The supplier will respond:

    “We follow industry best practice.”

    Ask which one.

    They will become irritated.

    This is progress.

    4. The Supplier Has Assured Itself

    One of the more delightful developments in enterprise technology is supplier-provided assurance.

    The supplier has reviewed the supplier’s design of the supplier’s product and concluded that the supplier recommends it.

    Excellent.

    A presentation will contain:

    PROVEN ARCHITECTURE

    ENTERPRISE GRADE

    SECURE BY DESIGN

    HIGHLY AVAILABLE

    SCALABLE

    There may be a Gartner logo.

    There will certainly be clouds.

    Your job is to ask:

    “Where is customer data stored?”

    The account manager will say:

    “In our secure cloud.”

    You ask:

    “Which country?”

    There will be a pause.

    “Within our global infrastructure.”

    This is not a country.

    Continue.

    5. Evidence Is Better Than Reassurance

    Programmes enjoy reassurance.

    “We’ve tested it.”

    “The supplier is confident.”

    “Security has been involved.”

    “Operations are comfortable.”

    “The business is happy.”

    None of these are evidence.

    Ask:

    Where are the test results?

    Where is the threat model?

    Where is the support model?

    Where is the capacity forecast?

    Where is the recovery test?

    Where is the data-flow diagram?

    Where is the licensing assessment?

    Who signed off the operational acceptance?

    At this point someone will accuse you of being “very detailed.”

    That is because evidence is offensive when reassurance was expected.

    6. The Red Flag Is Usually in the Footnote

    Executive summaries are optimistic.

    Detailed sections are cautious.

    Footnotes are where truth goes to hide.

    The first page may say:

    The proposed solution meets all strategic requirements.

    Page 73 may say:

    Note: current design does not provide automatic failover between sites.

    Page 106:

    Backup integration is outside current scope.

    Page 129:

    Existing identity platform is not formally supported.

    Page 151:

    Performance testing has not yet been scheduled.

    Page 162:

    Licensing position remains subject to vendor clarification.

    The executive summary will remain green.

    Your job is to make page 162 somebody’s problem.

    7. “Out of Scope” Is a Magical Phrase

    If something difficult cannot be solved, it can often be moved out of scope.

    Disaster recovery?

    Out of scope.

    Operational monitoring?

    Out of scope.

    Data migration reconciliation?

    Out of scope.

    Decommissioning?

    Out of scope.

    Security hardening?

    Phase Two.

    When enough things are out of scope, the project becomes very simple.

    It may no longer deliver a usable service.

    But the project is beautifully controlled.

    Solution Assurance should always ask:

    “Out of scope for whom?”

    If the answer is:

    “Operations will pick it up later.”

    You have discovered a landfill site.

    8. Phase Two Does Not Exist

    There is Phase One.

    Then there is production.

    Phase Two is a spiritual concept.

    It contains:

    technical debt,

    nice-to-have security controls,

    automation,

    performance optimisation,

    proper monitoring,

    full documentation,

    removal of temporary accounts,

    legacy decommissioning,

    and everything everyone promised would happen after go-live.

    Phase Two is funded from next year’s budget.

    Next year arrives.

    The programme is closed.

    A new transformation initiative begins.

    Phase Two becomes “legacy remediation.”

    Eventually it becomes somebody’s audit finding.

    9. RAG Status Is Applied Psychology

    Red means:

    Something is wrong.

    Amber means:

    Something is wrong but we are still discussing it.

    Green means:

    Nobody senior has asked the right question yet.

    There is also:

    Amber-Green.

    This is corporate synaesthesia.

    Amber-Green means:

    There are substantial concerns but the programme director has a board meeting.

    There is also:

    Green with commentary.

    This means:

    Please read the commentary.

    Nobody reads the commentary.

    10. Assurance Meetings Are Linguistic Combat

    A typical assurance meeting contains:

    the architect,

    the programme manager,

    the delivery lead,

    the supplier,

    security,

    operations,

    and somebody from PMO who has never spoken but is taking frighteningly good notes.

    You ask:

    “What happens if the database becomes unavailable?”

    Supplier:

    “The platform is highly resilient.”

    You:

    “How?”

    Supplier:

    “It uses a clustered architecture.”

    You:

    “What is the failover time?”

    Supplier:

    “It is designed for rapid recovery.”

    You:

    “What is the tested failover time?”

    Supplier:

    “We haven’t tested that scenario yet.”

    Programme Manager:

    “Is this really necessary for this stage?”

    You:

    “Yes.”

    Programme Manager:

    “Can we record it as an action?”

    This is how architecture becomes archaeology.

    11. The Action Log Is Where Risks Go to Hibernate

    Actions are useful.

    Until there are 147 of them.

    Every uncomfortable issue can be transformed into an action.

    Action 37: Confirm DR capability.

    Owner: Supplier.

    Due date: Friday.

    Friday arrives.

    Status:

    Open – awaiting supplier input.

    Next week:

    Open – supplier investigating.

    Next month:

    Open – to be addressed post go-live.

    Three months later:

    Closed – transferred to BAU.

    Nothing has actually happened.

    But the action is closed.

    Governance has achieved transcendence.

    12. “Accepted Risk” Requires Someone to Accept It

    Teams sometimes say:

    “The business has accepted the risk.”

    Ask:

    “Who?”

    A silence follows.

    “The business.”

    The business is not a person.

    The business does not have an email address.

    The business cannot attend court.

    The business cannot explain itself to an auditor.

    Risk acceptance requires an accountable individual with authority to accept the consequence.

    Find that person.

    Make them understand the risk.

    Get the acceptance recorded.

    You will be accused of bureaucracy.

    Ignore this.

    The same people will become intensely interested in documentation after the failure.

    13. Risk Language Must Describe Consequences

    Bad risk:

    There is a risk that the solution may not be resilient.

    Excellent.

    Meaningless.

    Better:

    Failure of the primary database node may cause complete service outage because automatic failover has not been implemented or tested. Recovery is dependent on manual intervention by the supplier. Estimated recovery time is unknown.

    Now management becomes interested.

    Especially the word unknown.

    Executives dislike unknown.

    This is useful.

    14. Assurance Is Not About Catching People Out

    Mostly.

    The objective is to expose uncertainty before uncertainty becomes outage.

    That requires uncomfortable questions.

    Not theatrical aggression.

    Do not enter a review saying:

    “This design is rubbish.”

    Ask:

    “What requirement led to this pattern?”

    Sometimes there is a good answer.

    Sometimes the answer is:

    “The vendor said so.”

    Sometimes:

    “We’ve always done it this way.”

    Sometimes:

    “We copied another project.”

    Sometimes nobody knows.

    The last one is surprisingly common.

    15. Architecture Debt Must Be Visible

    Every programme accumulates compromises.

    Temporary integration.

    Manual failover.

    Unsupported browser.

    Shared service account.

    Single region deployment.

    Missing monitoring.

    One firewall exception.

    One more firewall exception.

    A third firewall exception to make the first two work.

    Each seems reasonable alone.

    Together they form a service held together by hope.

    Assurance should aggregate these.

    A solution with twenty individually tolerable risks may be collectively intolerable.

    Programmes dislike this idea.

    They prefer risks in separate rows.

    Separate rows look smaller.

    16. Never Allow “Known Limitation” to Become “Normal”

    Known limitation is often corporate language for:

    We know it is broken.

    Examples:

    “The service requires restart every Sunday.”

    “The interface occasionally duplicates messages.”

    “Users must clear browser cache after upgrades.”

    “Failover can cause data inconsistency.”

    “Support requires local administrator access.”

    These may be genuine limitations.

    But they must have consequences, ownership and remediation.

    Otherwise three years later someone will say:

    “That’s just how it works.”

    This is how defects become culture.

    17. Test Evidence Is More Valuable Than Test Plans

    A test plan says:

    “We intend to test resilience.”

    Test evidence says:

    “We unplugged it and watched what happened.”

    Prefer the second.

    Ask for:

    load test results,

    failover evidence,

    restore evidence,

    penetration test findings,

    security scans,

    integration results,

    user acceptance outcomes.

    Never be impressed by:

    TESTING COMPLETE

    Ask:

    “What failed?”

    If the answer is:

    “Nothing.”

    Be suspicious.

    Either the system is extraordinary or the testing was decorative.

    18. If Nobody Tested Failure, Nobody Tested the System

    Successful transactions are pleasant.

    Failures are architecture.

    Disconnect the network.

    Kill the service.

    Expire the token.

    Remove the DNS record.

    Fill the disk.

    Lose the node.

    Corrupt the message.

    Throttle the API.

    Break authentication.

    Then observe.

    Systems reveal their real architecture when something goes wrong.

    19. Operational Acceptance Is Not “Ops Were Invited”

    Programs sometimes claim Operations has accepted a service because someone from Operations attended a meeting.

    That is not acceptance.

    Ask:

    Do they have monitoring?

    Runbooks?

    Access?

    Training?

    Escalation paths?

    Support contracts?

    Backup procedures?

    Recovery procedures?

    Capacity information?

    Known-error records?

    Service ownership?

    If not, Operations has not accepted the service.

    They have merely witnessed its birth.

    20. The Handover Document Is Usually Fiction

    A project handover document says:

    “BAU support will be provided by Infrastructure Services.”

    Infrastructure Services says:

    “Never heard of it.”

    Project:

    “They were on the distribution list.”

    Infrastructure:

    “So was Catering.”

    Project:

    “We assumed they were aware.”

    Operations:

    “We assumed you had a support contract.”

    Supplier:

    “Support contract?”

    Silence.

    The solution is now live.

    Congratulations.

    21. Security Exceptions Breed

    One exception is temporary.

    Two exceptions are pragmatic.

    Three exceptions are architecture.

    Assurance must track:

    what control is bypassed,

    why,

    who approved it,

    what compensating controls exist,

    when the exception expires.

    Otherwise the exception survives.

    After several years it becomes:

    legacy security model.

    People will then be afraid to remove it.

    22. Data Sovereignty Is Not Where the Salesman Lives

    Ask where data resides.

    The vendor says:

    “UK hosted.”

    Ask:

    Backups?

    Logs?

    Telemetry?

    Support access?

    Disaster recovery?

    Sub-processors?

    AI services?

    Analytics?

    Suddenly the United Kingdom becomes geographically flexible.

    Assurance exists partly to continue asking after the first comforting answer.

    23. “Encrypted” Is the Beginning of the Question

    Encrypted where?

    At rest?

    In transit?

    Client-side?

    Server-side?

    Which algorithm?

    Who holds the keys?

    Can the provider decrypt it?

    Can administrators?

    How are keys rotated?

    What happens during recovery?

    If the vendor says:

    “AES-256.”

    Do not applaud.

    AES-256 is not an architecture.

    It is a cipher.

    24. Backup Is Not Resilience

    A backup protects data.

    It does not automatically protect:

    availability,

    configuration,

    identity,

    network connectivity,

    integration state,

    DNS,

    certificates,

    secrets,

    or the ability of anyone to remember how to restore the bloody thing.

    Assurance should ask for recovery, not backup.

    “How quickly can you restore the service from nothing?”

    Watch confidence decrease.

    This is healthy.

    25. DR Documentation Is Often Fantasy Literature

    The DR plan may say:

    In the event of primary site loss, service will fail over to secondary site.

    Ask:

    Who performs the failover?

    “How?”

    “Using the DR process.”

    “Where is that?”

    “In the DR document.”

    You are already reading the DR document.

    This is recursive resilience.

    Continue until someone admits Gary knows how.

    26. Capacity Is Not “Scalable”

    Supplier:

    “The platform scales automatically.”

    You:

    “To what?”

    Supplier:

    “As demand increases.”

    You:

    “What is the tested maximum?”

    Supplier:

    “That depends on configuration.”

    You:

    “What configuration are we buying?”

    Supplier:

    “We can confirm that during implementation.”

    Programme:

    “Can we move on?”

    No.

    We cannot.

    27. Performance Requirements Need Numbers

    “Fast.”

    No.

    “Responsive.”

    No.

    “Near real time.”

    Absolutely not.

    Use:

    95th percentile response under two seconds.

    10,000 concurrent users.

    500 transactions per second.

    Batch completion before 06:00.

    Data propagation within 30 seconds.

    Numbers can be tested.

    Adjectives can only be discussed.

    28. Monitoring Is Not a Dashboard Nobody Watches

    A colourful dashboard is not monitoring.

    Ask:

    Who receives alerts?

    What thresholds exist?

    What constitutes service degradation?

    Who responds?

    How quickly?

    Are alerts tested?

    Are dependencies monitored?

    Can users be affected while every infrastructure metric remains green?

    The answer to the last question is usually yes.

    Infrastructure can be perfectly healthy while the application is utterly fucked.

    This is why service monitoring exists.

    In theory.

    29. Observability Is Not Logging Everything

    Modern systems can produce terrifying quantities of logs.

    This is not observability.

    Observability means being able to answer:

    What happened?

    Where?

    When?

    To whom?

    Why?

    Across which components?

    A petabyte of JSON nobody can correlate is merely expensive confusion.

    30. Compliance Is Not Security

    A system may pass an audit and still be insecure.

    A system may be secure and still fail compliance.

    These overlap.

    They are not identical.

    Checkbox security produces magnificent evidence packs.

    Attackers do not generally read them.

    31. The Penetration Test Is Not an Exorcism

    A penetration test does not bless the system.

    It tests a defined scope at a point in time.

    Ask:

    What was excluded?

    Was authentication tested?

    APIs?

    Internal interfaces?

    Cloud configuration?

    Privilege escalation?

    Mobile clients?

    Infrastructure?

    Was the production configuration actually tested?

    The executive summary will say:

    No critical findings.

    Page twelve may contain eight High findings.

    Read page twelve.

    32. “Low Risk” Findings Can Combine Into a High Risk System

    Weak password policy.

    Verbose error messages.

    Excessive permissions.

    Unrestricted outbound connectivity.

    Poor logging.

    Individually low or medium.

    Together:

    Excellent afternoon for an attacker.

    Assurance must think in systems.

    Risk registers often do not.

    33. Dependencies Are Where Assurance Earns Its Keep

    The application may be resilient.

    But it depends on:

    DNS,

    identity,

    network,

    API gateway,

    certificate authority,

    message broker,

    database,

    storage,

    third-party payment gateway,

    and an ancient file transfer server in Swindon.

    Ask what happens when each fails.

    Someone will say:

    “That is outside our solution boundary.”

    Failure does not respect solution boundaries.

    34. The Boundary Diagram Is a Negotiation

    Projects draw solution boundaries partly to define ownership.

    Unfortunately, the customer experience does not care.

    If your beautifully assured application cannot function because the corporate proxy is unavailable, the service is unavailable.

    Users will not say:

    “Fortunately, the application component remained compliant with its architecture.”

    They will say:

    “It doesn’t fucking work.”

    Assure the service.

    Not just the boxes.

    35. Third Parties Are First-Class Risks

    Vendor:

    “We use a specialist third party for that.”

    Assurance:

    “Who?”

    Vendor:

    “That information is commercially sensitive.”

    Assurance:

    “They process our data.”

    Vendor:

    “We can provide details under NDA.”

    Good.

    Continue.

    Subcontracting does not outsource accountability.

    It merely lengthens the incident bridge.

    36. Exit Strategy Is Architecture

    Every supplier relationship ends.

    Eventually.

    Ask:

    How do we retrieve data?

    In what format?

    How long does extraction take?

    What does it cost?

    Can another provider consume it?

    What happens to backups?

    When is data deleted?

    How do we verify deletion?

    What happens if the supplier becomes insolvent?

    Procurement may consider these gloomy questions.

    They are.

    So is divorce law.

    Still useful.

    37. Licensing Assurance Exists Because Lawyers Enjoy Ambiguity

    Technical teams think software licensing is about software.

    It is actually about contractual nouns.

    Installed.

    Used.

    Accessed.

    Processor.

    Core.

    Named user.

    Authorised user.

    Indirect use.

    Backup.

    Failover.

    Test.

    Development.

    Virtualisation.

    Cloud mobility.

    Multiplexing.

    Ask whether the architecture changes licence exposure.

    If nobody knows, record the uncertainty.

    Do not allow:

    “The account manager said it was fine.”

    The account manager will not attend the audit.

    38. Cost Assurance Must Include Success

    Projects estimate cost at average load.

    Success changes this.

    More users.

    More storage.

    More API traffic.

    More logs.

    More backups.

    More egress.

    More licences.

    Ask:

    “What does this cost if adoption is twice forecast?”

    If the answer is:

    “That would be a good problem to have.”

    You have found someone who does not pay cloud bills.

    39. Assumptions Are Risks Wearing Fake Moustaches

    Designs contain assumptions.

    “Existing WAN has sufficient capacity.”

    “Users have modern browsers.”

    “Partner API supports required volumes.”

    “Directory contains accurate attributes.”

    “Legacy system will remain available.”

    “We assume 20% annual growth.”

    Assumptions should be validated.

    Otherwise they are risks disguised grammatically.

    40. “To Be Confirmed” Has an Expiry Date

    TBC is acceptable early.

    Later, it becomes dangerous.

    At design review:

    TBC.

    At build:

    TBC.

    At test:

    TBC.

    At go-live:

    TBC.

    In the incident report:

    Root cause.

    Every TBC should have:

    an owner,

    a due date,

    and consequences if unresolved.

    Otherwise you are manufacturing uncertainty professionally.

    41. Decision Ownership Matters

    Who decided to accept single-region hosting?

    Who decided not to implement automated recovery?

    Who approved the unsupported integration?

    Who accepted the licence risk?

    The answer cannot be:

    “The programme.”

    Programs do not go to disciplinary hearings.

    People do.

    Architecture decisions require named ownership.

    This tends to improve decision quality dramatically.

    42. Escalation Is Not Failure

    Assurance practitioners sometimes avoid escalation because they do not want to appear obstructive.

    This is cowardice wearing stakeholder-management clothing.

    If a material risk exceeds your authority, escalate it.

    Calmly.

    With evidence.

    Without drama.

    Then let the accountable person decide.

    Your job is not to win.

    Your job is to ensure the decision is conscious.

    43. A Waiver Is Not a Magic Spell

    Sometimes a programme requests an architectural waiver.

    Fine.

    A waiver should state:

    what standard is being waived,

    why,

    risk created,

    compensating controls,

    owner,

    expiry.

    A permanent waiver is not a waiver.

    It is a new standard nobody has admitted exists.

    44. Mature Assurance Knows When to Stop

    Not every system requires military-grade resilience.

    Not every application needs active-active deployment across continents.

    Not every dataset needs hardware-backed encryption keys rotated hourly.

    Assurance must be proportionate.

    Risk depends on consequence.

    The lunch-menu application can occasionally fail.

    The air-traffic control system should perhaps have stronger aspirations.

    Apply judgement.

    Otherwise assurance itself becomes the risk.

    45. The Assurance Practitioner Must Understand Delivery

    A reviewer who has never built anything is dangerous.

    They may demand:

    perfect documentation,

    zero technical debt,

    complete automation,

    full resilience,

    maximum security,

    infinite scalability,

    and delivery by Friday.

    Architecture is trade-off.

    Assurance must understand trade-off.

    Ask whether risk is conscious and proportionate.

    Do not demand utopia.

    Utopia is not supportable.

    46. “Industry Best Practice” Is Often Consultancy Incense

    Whenever someone invokes best practice, ask:

    For this context?

    For this scale?

    For this threat model?

    For this regulatory environment?

    For this operating model?

    A multinational bank and a village museum do not necessarily require identical controls.

    If your assurance framework says they do, the framework is the thing requiring assurance.

    47. Templates Are Useful Until They Replace Thinking

    Assurance templates provide consistency.

    Good.

    But the template does not know what is important.

    A reviewer may spend twenty minutes checking whether every heading is populated while missing the fact that the system has no backup.

    This is called compliance theatre.

    Never confuse completeness of form with completeness of thought.

    48. Architecture Boards Attract PowerPoint

    A board pack may contain:

    executive summary,

    strategic alignment,

    business outcomes,

    capability mapping,

    technology principles,

    risk summary,

    implementation roadmap.

    Excellent.

    Ask:

    “What port does it use?”

    Nobody knows.

    This is not because ports are strategically important.

    It is because detail reveals whether anybody has actually designed anything.

    Move between levels.

    That is the job.

    49. The Best Assurance Question Is Often “Show Me”

    “We have backups.”

    Show me a restore.

    “We monitor it.”

    Show me an alert.

    “We tested failover.”

    Show me the evidence.

    “Operations accepted it.”

    Show me the acceptance.

    “The vendor supports this.”

    Show me where.

    “Security approved it.”

    Show me the decision.

    “Licensing is covered.”

    Show me the entitlement.

    This phrase eliminates approximately seventy percent of enterprise bullshit.

    Use responsibly.

    50. Assurance Must Survive Executive Pressure

    Someone important will eventually say:

    “Can you just sign this off?”

    No.

    You can review it.

    You can assure it.

    You can identify conditions.

    You can record risks.

    You cannot transform uncertainty into certainty because a meeting starts at 14:00.

    If necessary say:

    “I can provide assurance based on the evidence available.”

    This is polite.

    It also creates a clear boundary around reality.

    51. Beware the Urgent Executive Exception

    There is always one.

    “We need to bypass the normal process.”

    Why?

    “Business-critical.”

    Everything is business-critical shortly before a board meeting.

    Urgency may justify accelerated assurance.

    It does not justify no assurance.

    Fast decisions need clearer risk statements, not fewer.

    52. The Incident Will Reopen Every Argument

    After an outage, people become historians.

    Someone will say:

    “Nobody could have predicted this.”

    Check your assurance report.

    Often, someone did.

    Another will say:

    “This was an unforeseeable dependency.”

    Check the architecture review.

    It may be listed.

    A third will say:

    “We understood the risk.”

    Ask for acceptance.

    Silence.

    This is why documentation matters.

    Not for blame.

    For organisational memory.

    Blame is merely a side effect.

    53. Lessons Learned Are Usually Lessons Observed

    The post-incident review produces:

    Improve documentation.

    Engage stakeholders earlier.

    Strengthen testing.

    Clarify ownership.

    Review monitoring.

    These lessons have appeared in every enterprise incident review since approximately 1987.

    A lesson is not learned because it is written.

    It is learned when behaviour changes.

    Otherwise it is merely rediscovered wisdom.

    54. Assurance Findings Need Closure Criteria

    Finding:

    “Improve monitoring.”

    Impossible to close meaningfully.

    Better:

    “Implement synthetic transaction monitoring for customer login and payment workflows, with alerts routed to 24×7 support and tested before production release.”

    Now closure can be evidenced.

    Specificity is the enemy of ceremonial governance.

    55. Do Not Let the Programme Mark Its Own Homework

    Programme:

    “We have resolved Finding 12.”

    Assurance:

    “How?”

    Programme:

    “We discussed it.”

    No.

    Resolution requires evidence.

    Otherwise the student has written:

    Corrected

    in the margin of their own exam paper.

    56. Some Risks Should Stop Go-Live

    This will upset people.

    Good.

    Examples may include:

    known exploitable security defects,

    no viable recovery capability for a critical service,

    unresolved data-loss risk,

    unsupported production configuration,

    absence of required regulatory controls,

    major capacity failure under expected load.

    A go-live date is not a law of physics.

    Sometimes the correct assurance outcome is:

    “No.”

    This is rare.

    It should remain available.

    Otherwise assurance is merely decorative.

    57. “Conditional Go-Live” Means Conditions

    Do not approve go-live subject to ten conditions which cannot realistically be completed after go-live.

    That is not conditional approval.

    That is denial with poor emotional resilience.

    If the condition matters before production, require it before production.

    If it can genuinely follow, assign:

    owner,

    date,

    risk,

    escalation.

    Words must mean things.

    This principle is surprisingly controversial.

    58. Assurance Should Reduce Surprise

    Perfect systems do not exist.

    Incidents will happen.

    The objective is not zero failure.

    It is fewer stupid surprises.

    You should not discover after go-live that:

    the backup never worked,

    the supplier does not provide 24×7 support,

    the application cannot run in DR,

    the licence excludes virtualisation,

    the logs contain personal data,

    the database has a 2TB limit,

    the certificate renewal is manual,

    or the only administrator is on maternity leave.

    These are not black swans.

    They are pigeons standing directly in front of you.

    59. The Highest Form of Assurance Is Boring Production

    No P1 incidents.

    Predictable patching.

    Tested recovery.

    Clear ownership.

    Known capacity.

    Controlled change.

    Understandable costs.

    Useful monitoring.

    Boring.

    Architects sometimes dislike boring systems.

    Operations loves them.

    Customers rarely complain that their transaction completed without architectural excitement.

    Boring is underrated.

    60. The Final Tao

    The novice believes Solution Assurance exists to approve solutions.

    The experienced practitioner knows it exists to expose decisions.

    The master understands that the organisation will sometimes make the wrong decision anyway.

    Your role is therefore not omnipotence.

    It is clarity.

    Make assumptions visible.

    Make dependencies visible.

    Make consequences visible.

    Make ownership visible.

    Ask for evidence.

    Challenge optimism.

    Separate confidence from fact.

    Do not allow green status to erase red engineering.

    Do not allow urgency to repeal physics.

    Do not allow governance to replace judgement.

    And above all, never write:

    “No significant architectural risks identified.”

    unless you have looked very, very hard.

    Because six months later, at 02:17 on a Sunday morning, when production is down, the backup is corrupt, DNS is pointing at the wrong data centre, the supplier’s support desk is closed, the certificate expired yesterday and nobody can remember who owns the service, somebody will find your assurance report.

    They will scroll to the final page.

    They will read your name.

    And they will ask the oldest question in enterprise architecture:

    “Who the fuck signed this off?”

    At that moment, enlightenment is achieved.

    Usually by someone else.