---
title: "Narrowing the jack mackerel CMP comparison"
subtitle: "Working paper: best-relative scorecards, management trade-offs and matched simulation trajectories"
date: 2026-09-10
lang: en
format:
  html:
    toc: true
    toc-depth: 2
    number-sections: true
    embed-resources: true
    lightbox: true
    code-fold: true
    theme: cosmo
    fig-width: 10
---

[Back to the jack mackerel wiki](../../mse.html) · [All CMPs and fixed catch: best-relative scorecard](scorecard-all.html) · [Retained CMPs and fixed catch: best-relative scorecard](scorecard-shortlist.html)

::: {.callout-note title="Purpose"}
This working paper supports discussion of CMP choices using the saved 500-simulation results. Model settings and tuning are unchanged; selection of a CMP remains open. This revision focuses on the recruitment-crash and recruitment-cycle OMs and adds the supplied fixed-catch reference comparison. The three-CMP presentation is a discussion set; the revised advice indicator requires the shortlist to be reviewed. Robustness results remain separate. The full set remains available in the scorecard and downloadable data.
:::

## Findings for discussion

Three CMPs provide a compact view of the main management choices: **HS−20 (MP43), HS−30 (MP47), and PR−20 (MP44)**. HS−20 represents the lower-variability hockey-stick alternatives with a 20% advice-reduction limit. HS−30 retains the contrast offered by allowing larger reductions. PR−20 retains the higher-catch power-ramp alternative and its greater interannual variability.

The three-CMP display is retained for continuity, but the revised advice-reduction statistic changes the earlier screening result. HS+20, HSsym, PR+20 and PRsym now present trade-offs with their listed comparators; their former exclusion is not supported by the updated five-indicator screen. Review all eight CMPs before confirming a shortlist. PR−30 remains poorer than HS−30 within the illustrative tolerances. Agreement on acceptable risk remains a separate step.

## Data, periods and indicator definitions

The results cover eight candidate management procedures (CMPs), each with 500 simulations under the reference operating model (OM) and nine alternative OMs used to test performance under different assumptions (robustness tests). Two-stock OMs are shown by biological component; their catches and biomasses are reported separately. Advice reductions use all annual advice comparisons across all iterations (advice years **2025–2049**, applied **2026–2050**). The near-term plots and biological probabilities use **2026–2035**. Long-term performance calculations use **2041–2050**, except the inherited mean-catch-reduction column available as an optional scorecard indicator, which uses **2026–2050**. The worm plots span **2025–2050**.

For catch, interannual catch change (IACC), and SSB/SSBMSY, first calculate each simulation's mean over 2041–2050. Trade-off points then show the median across 500 simulations; horizontal and vertical bars show the 25th–75th percentiles. The bars contain the middle half of simulation-specific values. Uncertainty in the estimated median would require a separate calculation. IACC is the absolute percentage change from the previous year's realized catch. The 2041 change therefore uses catch in 2040. The biomass ratio in the trade-off plots uses the stored **static SSBMSY** denominator; it is distinct from the dynamic-B0 risk indicator.

The low-biomass indicator is the proportion of simulation-years with **SSB <8% of the matching year's dynamic B0**, the previous Blim. Dynamic B0 is the saved projected unfished SSB for the matching simulation. Each year below the threshold counts, including consecutive years. Catch below 270 kt is also a pooled simulation-year frequency.

The **advice-reduction frequency** (`CatchDrop20`) is the proportion of valid annual HCR-advice comparisons with $A_t < 0.81 A_{t-1}$, pooled over all years and 500 iterations. Advice is recorded in 2025–2049 and applied in 2026–2050. The first advice is compared with the saved HCR initial advice, or the constant advice target for fixed catch. This gives 25 comparisons per iteration (12,500 when complete). Exact 19% reductions and increases do not count; 20% reductions count. There are no biomass or near-20% exclusions. Invalid values and nonpositive previous advice are excluded from the denominator and reported in the audit. This uses advice before implementation. Two-stock OMs have advice for the managed Southern stock; the same run-level advice statistic appears in both component views.

[All OM/CMP advice counts and denominators](catchdrop19-advice-summary.csv) · [Annual advice events and checks](candidate_quilt_catch_cut_diagnostics_reference.csv) and [counts and eligible years per simulation](candidate_quilt_event_counts_reference.csv) retain the audit trail.

## Screening the CMP set

Five management-facing indicators are used for this screen: median catch, IACC, low-catch probability, SSB-below-previous-Blim frequency, and the advice-reduction frequency. All eleven indicators remain available in the scorecards; the shortlist uses the five indicators listed here.

The following absolute differences are treated as small for **forming this discussion shortlist**. These thresholds remain open for review. The former tolerance of 0.1 event per ten-year window is expressed as 1 percentage point for the revised advice frequency; its use with the new definition also requires review:

| Indicator | Screening tolerance | Preferred direction |
|---|---:|---|
| Median catch | 40 thousand t | Higher |
| IACC | 0.6 percentage point | Lower |
| Probability of catch below 270 kt | 1 percentage point | Lower |
| Frequency below previous Blim | 1 percentage point | Lower |
| Advice reductions >19% | 1 percentage point | Lower |

A **close alternative** differs from its retained comparator by no more than every listed tolerance. A **poorer performer within these tolerances** has a retained comparator that is no worse beyond any tolerance and better beyond at least one. Every omitted CMP is compared directly with a retained CMP; there is no chain of eliminations. Choosing HS−20 and PR−20 as the representatives preserves the focal minus-20 variants for discussion. This presentation choice does not establish that the omitted CMPs can be discarded under the revised indicator; the overall score is displayed separately.

{{< include screening.md >}}

HS+20, HSsym, PR+20 and PRsym have no advice reductions greater than 19% in the reference runs, whereas their retained comparators do. These differences exceed the illustrative one-percentage-point tolerance and create trade-offs, so the earlier close-alternative or poorer labels for those four CMPs no longer apply. PR−30 remains poorer than HS−30 under these five indicators and tolerances. All comparisons remain open to revision for other OMs, indicators or stakeholder preferences.

### Full reference comparison

The table retains the omitted CMPs so the screen can be challenged. C is thousand tonnes; IACC is percent; PC270 and SSBbelow8dB0 are displayed as percentages; CatchDrop20 is displayed as the percentage of all valid consecutive advice comparisons with reductions greater than 19%.

{{< include all-cmp-values.md >}}

[Screening decisions](screening.csv) · [Screening tolerances](screening-tolerances.csv) · [All eleven quilt indicators](all-cmp-quilt-values.csv)

## Scorecards: best 100, others relative to it

Use the [all-CMP scorecard](scorecard-all.html) to review the original set and the [retained-set scorecard](scorecard-shortlist.html) for the retained-set comparison and fixed-catch reference comparison. Both open with **best-relative scaling** and equal weights on the available screening indicators. Use **Choose an OM** to select **reference (OM11), recruitment crash (OM11_2), or recruitment cycle (OM11_3)**. Scores are recomputed separately within each OM. The full nine-OM robustness plots and data remain available in the expandable appendix below.

The fixed-catch benchmark is available in the reference selection only. The supplied fixed-catch results cover the reference OM. When fixed catch is selected, the two vulnerable-biomass indicators are unavailable pending extraction from the supplied run; deselect it to restore those choices. Each scorecard displays only indicators shared by the selected CMPs. The fixed-catch snapshot uses its own stored reference points and reconstructed dynamic B0, preserving its reported tuning convention. Users can change CMPs, indicators, scaling and weights. The overall score helps explore preferences. The revised five-indicator comparison requires the earlier shortlist to be reviewed.

For a higher-is-better indicator, the score is $100V_i/\max_j(V_j)$. For a lower-is-better indicator, it is $100\min_j(V_j)/V_i$. The best value therefore scores 100. Scores are recalculated for the selected CMP set, so each selection has its own reference for comparison.

**Zero-best exception:** when the best value for a lower-is-better indicator is zero, zero-valued CMPs score 100 and positive-valued CMPs score zero. This rule is necessary because the ordinary ratio cannot distinguish positive values when its numerator is zero. In the full reference set, several CMPs and fixed catch have zero advice reductions greater than 19%, while HS−20, HS−30, PR−20 and PR−30 have positive frequencies. This gives the latter zero scores for this indicator when a zero-valued comparator is selected. Inspect the raw frequencies and compare with range scaling before assigning weight. Related stock or stability indicators can overlap even when given equal weights.

The [existing general best-relative explorer](https://sprfmo.github.io/jmMSE26/application/scorecard-best-relative.html) and [original range-scaled explorer](https://sprfmo.github.io/jmMSE26/application/scorecard.html) provide the broader application context. The two companion scorecards above embed this paper's exact data snapshot for reproducibility.

## Near-term performance: 2026–2035

The near-term plots use the first ten projection years, **2026–2035**, for the reference OM, recruitment crash (OM11_2) and recruitment cycle (OM11_3). Each point is the median across 500 iteration-specific period means. The trade-off bars show the 25th–75th percentiles in both directions; they describe the distribution of simulation outcomes.

![Near-term catch and interannual catch variability for the three focused OMs.](near-term-IACC.png){#fig-near-iacc fig-alt="Three panels show 2026 to 2035 median catch horizontally and median interannual catch variability vertically, with horizontal and vertical interquartile ranges for each CMP."}

![Near-term catch and spawning-biomass status using the stored static SSBMSY denominator.](near-term-SBMSY.png){#fig-near-ssb fig-alt="Three panels compare 2026 to 2035 catch with SSB divided by static SSBMSY. Points are medians of 500 iteration means, with interquartile bars on both axes."}

![Near-term Kobe status for the reference, recruitment-crash and recruitment-cycle OMs.](near-term-kobe.png){#fig-near-kobe fig-alt="Kobe panels for 2026 to 2035. Spawning biomass relative to dynamic SSBMSY is horizontal and fishing mortality relative to FMSY is vertical. Labelled CMP points and reference lines at one distinguish the four status regions."}

The Kobe biomass denominator is **dynamic SSBMSY**, calculated as the stored SSBMSY/SSB0 ratio times that year's and iteration's unfished SSB. It matches the dynamic convention used for tuning and the existing reference Kobe figure. The catch–biomass trade-off instead retains the **static SSBMSY** convention used in the long-term plots. A median Kobe point is a summary of status, not a probability of meeting a target.

The following probabilities pool all **5,000 near-term year–iteration outcomes per CMP/OM**. Green means SSB/dynamic SSBMSY ≥1 and F/FMSY ≤1; red means both SSB/dynamic SSBMSY <1 and F/FMSY >1. SSB below the previous Blim uses 8% of dynamic B0. These period-specific biological probabilities are separate from `CatchDrop20`, which continues to use all 25 projection years and 500 iterations.

{{< include near-term-risk.md >}}

[Near-term trade-off data](near-term-tradeoffs.csv) · [Kobe point data](near-term-kobe.csv) · [Near-term biological probabilities](near-term-risk.csv) · [Input checks](near-term-provenance.txt)

## Reference-OM trade-offs after screening

![Catch and stability among retained CMPs. Horizontal and vertical bars are interquartile ranges across the 500 simulation-specific 2041–2050 means.](reference-IACC.png){#fig-reference-stability fig-alt="Three retained CMPs plotted with median catch on the horizontal axis and median interannual catch change on the vertical axis. Both directions have interquartile bars; colours and point shapes distinguish CMPs."}

![Catch and spawning-biomass status among retained CMPs, using the stored static SSBMSY reference point.](reference-SBMSY.png){#fig-reference-biomass fig-alt="Median catch versus median SSB divided by static SSBMSY for HS minus 20, HS minus 30 and PR minus 20, with horizontal and vertical interquartile ranges."}

The reference plots expose the cost of concentrating on one outcome. The preferred CMP can change with the priority given to catch stability, catch level or low-biomass risk. Simulation ranges can overlap while differences between rules remain within individual simulations. The shortlist should be revisited if a different time horizon, screening tolerance or biological constraint is agreed.

## Focused robustness: recruitment crash and recruitment cycle

The focused comparison retains **HS−20 (MP43), HS−30 (MP47) and PR−20 (MP44)** under **OM11_2 (h1_0.16_lowrec, recruitment crash)** and **OM11_3 (h1_0.16_cycle, recruitment cycle)**. Both use the saved single-stock simulations and existing tuning. The reference OM is included for orientation. OM11_3 replaces OM21 in the focused comparisons and demonstrations.

![Long-term catch and interannual variability for the reference, recruitment-crash and recruitment-cycle OMs.](focused-robustness-IACC.png){#fig-focus-iacc fig-alt="Three single-stock panels compare selected CMPs using 2041 to 2050 median catch and catch variability with horizontal and vertical interquartile ranges."}

![Long-term catch and static-reference SSB status for the same three OMs.](focused-robustness-SBMSY.png){#fig-focus-ssb fig-alt="Three panels compare 2041 to 2050 catch and SSB divided by static SSBMSY, with interquartile ranges on both axes."}

The biological risk table uses **dynamic SSBMSY** for green status and **8% of dynamic B0** for the previous Blim. It pools the **5,000 year–iteration outcomes per CMP/OM over 2041–2050**. OMs remain separate.

{{< include focused-robustness-risk.md >}}

Under recruitment crash, PR−20 has the highest long-term dynamic-green probability (54.5%) and the lowest below-previous-Blim frequency (5.0%) of these three. Under recruitment cycle, PR−20 also has the highest green probability (51.5%) and the lowest below-Blim frequency (7.7%). HS−20 and HS−30 have nearly equal cycle green probabilities (44.1% and 44.2%), but HS−30 has a lower below-Blim frequency (10.6% versus 13.2%). Catch level and variability should be considered alongside these biological outcomes.

![Fifteen matched SSB and catch trajectories under recruitment crash.](worms-15-om11_2.png){#fig-worm-rec-crash fig-alt="Six panels show the same 15 iteration IDs for SSB and catch under the three selected CMPs in the recruitment-crash OM. Low outcomes remain included."}

![Fifteen matched SSB and catch trajectories under recruitment cycle, OM11_3.](worms-15-om11_3.png){#fig-worm-rec-cycle fig-alt="Six panels show SSB and catch for the same 15 iteration IDs under the three selected CMPs in recruitment-cycle OM11_3. Vertical scales are shared within each row."}

[Focused trade-off data](focused-robustness-tradeoffs.csv) · [Dynamic risk results](focused-robustness-risk.csv) · [Robustness trajectories](focused-robustness-worms.csv) · [Focused input record](focused-robustness-provenance.txt)

<details>
<summary>Full robustness plots and interpretation across all nine OMs</summary>


The alternative OMs provide a separate check of the reference shortlist. The following plots keep each OM and stock component separate. Panel scales vary to make within-panel comparisons legible; compare numerical axes before comparing distances across panels. No probabilities or weights have been assigned to the OMs.

![Catch–stability trade-offs for the five single-stock robustness OMs.](robustness-single-IACC.png){#fig-single-stability fig-alt="Five panels show retained CMP catch and interannual catch-change medians and interquartile ranges for alternative selectivity, low recruitment, cyclic recruitment, higher steepness and the alternative assessment model."}

![Catch–biomass trade-offs for the five single-stock robustness OMs.](robustness-single-SBMSY.png){#fig-single-biomass fig-alt="Five separate single-stock robustness panels compare catch and static-reference SSB ratios for the retained CMPs, with interquartile bars on both axes."}

![Catch–stability trade-offs for each biological component of the four two-stock robustness OMs.](robustness-two-IACC.png){#fig-two-stability fig-alt="Eight panels separate the two biological components of four two-stock OMs. Each displays retained CMP catch and variability with interquartile ranges."}

![Catch–biomass trade-offs for each biological component of the four two-stock robustness OMs.](robustness-two-SBMSY.png){#fig-two-biomass fig-alt="Eight stock-component panels compare catch and static-reference spawning-biomass ratios under two-stock, movement, steepness and assessment-model alternatives."}

Among the retained CMPs, PR−20 has the highest median catch in the reference
and alternative-selectivity OMs, while one of the hockey-stick CMPs has the
higher catch in the other non-collapsed panels. PR−20 often retains a higher
SSB/static-SSBMSY ratio under robustness OMs. These contrasts reinforce the
need to consider yield and biomass jointly rather than extrapolating the
reference-OM ranking. The northern component of h2_1.14 has essentially zero
catch under all three CMPs; that panel is labelled as having too little catch to support a
useful ranking. Its raw values remain in the downloadable data.

All eight CMPs' OM-specific summaries are retained in the [complete trade-off data](tradeoff-summary-all-cmps.csv). An omitted CMP should be reinstated if the working group identifies a useful advantage under an alternative OM. The preferred CMP across OMs remains open. The shortlist is based on practical comparison thresholds; formal statistical testing would be a separate analysis.


</details>

## Fifteen matched simulation paths {#fifteen-matched-simulation-worms}

The 15 iteration IDs are reused from the earlier paired reference figure: **37, 65, 110, 151, 158, 165, 177, 186, 246, 327, 344, 348, 367, 397 and 498**. The same IDs and colours are used in all six panels. The original selection is retained, including poor outcomes. Matching IDs align the underlying parameter draws. Random observation errors can still differ between CMP runs that use different seeds.

![SSB and catch trajectories for the same 15 simulations under each retained CMP.](worms-15.png){#fig-worms fig-alt="Six panels show SSB above and catch below for HS minus 20, HS minus 30 and PR minus 20. Fifteen individually coloured simulations are matched across all panels; all trajectories, including low outcomes, are retained."}

These plots show the timing and persistence of changes in 15 of the 500 simulations. Outcome probabilities and uncertainty ranges are calculated from the full set. The full-distribution trade-off summaries and low-biomass frequencies remain the basis for comparing typical performance and risk.

[Iteration IDs](worm-iterations.csv) · [SSB and catch trajectories](worm-trajectories.csv)

## Fixed-catch reference comparison {#fixed-catch-reference-benchmark}

The supplied `tunfixc.rds` is an `FLmse` reference-OM result with **500
simulations**, projection years 2025–2050, and `fixedC.hcr` advice of **1,525
thousand tonnes per year**. It is intended to be tuned to the same 60% dynamic
green objective as the other CMPs. Reconstructing its unfished trajectory with
the repository's zero-catch projection, using the supplied run's own initial
state, gives **P(green) = 60.3% over 2041–2050**. This is consistent with the
reported target. This checks the achieved probability. Reviewing the full tuning history
would require the original tuning record and stopping tolerance, which are
absent from the supplied result. Using static BMSY instead gives 57.8%.

{{< include fixed-catch-summary.md >}}

Advice is fixed while realized catch varies. Recorded
implementation-adjusted targets vary from about **1,517 to 1,593 thousand
tonnes**; the mean catch within each simulation therefore varies, as do
realized annual changes. The supplied run includes implementation adjustments,
whereas the current reference tuning call for the three CMPs omitted
the random implementation-error model.

The initial numbers-at-age, mortality, maturity, weights and stock–recruit
parameters match the current reference OM. However, the projected recruitment
deviations and stored MSY reference points differ. For example,
median stored FMSY is about **0.298** in the fixed-catch run and **0.334** in
the current reference bundle. The scorecard includes it as a reference
comparison using its own saved reference points. Isolating the effect of
the harvest control rule (HCR) would require a rerun with matching
recruitment, reference points and implementation assumptions.

![SSB trajectories for 15 unique simulations from the supplied fixed-catch reference run.](worms-15-fixed-catch.png){#fig-worm-fixed fig-alt="One panel shows SSB for 15 distinct simulations of the supplied fixed-catch reference run from 2025 to 2050, including trajectories that fall sharply."}

The same 15 display IDs help orient the reader; recruitment paths differ
from those in the current CMP runs. The available fixed-catch results cover
the reference OM, so this option appears only in the reference scorecard.

[Fixed-catch summary](fixed-catch-summary.csv) ·
[Stored reference-point comparison](fixed-catch-reference-point-check.csv) ·
[Recorded implementation targets](fixed-catch-implementation-targets.csv) ·
[Input files and reconstruction record](robustness-fixedcatch-provenance.txt)

## Decisions this paper can support

The companion [draft decision guide](cmp-decision-guide.html) uses these results to propose advice, a manager decision checklist and the work needed before an adoption recommendation. It is a discussion draft for review by the Scientific Committee, managers and Parties.

1. Review the explicit screening tolerances and decide whether the three-CMP presentation set is sufficient for the next discussion.
2. Set biological acceptability criteria separately from best-relative scoring; a relatively high score can coexist with unacceptable risk.
3. Examine robustness panels and the full-set data before treating an omitted CMP as dispensable.
4. Agree which indicators and weights represent the objectives, paying particular attention to zero-best risk indicators and overlapping measures.
5. Retain projection-failure diagnostics and advice-versus-realized-catch checks as outstanding validation work. The advice statistic measures decisions before implementation; failed projections and their effects on realized-catch indicators still require investigation.

## Source files and checks {#reproducibility-and-scope-of-validation}

The base script `R/build_cmp_screening_paper.R` reads the saved reference and robustness performance tables from `output/candidate-performance-500/`, and the reference runs from `jmMSE-500-refine/model/tune/refine_500_from_100/runs.rds`. The saved CMP simulations are unchanged. A zero-catch counterfactual was reconstructed only to check dynamic reference points for the fixed-catch handoff. The current quilt supplies the advice-reduction probability and dynamic-B0 risk values. The extension `R/extend_cmp_paper_robustness_fixedcatch.R` reads the OM11_2 and OM11_3 robustness checkpoints and inspects the unchanged Downloads input. `R/build_cmp_om_scorecards.R`, `R/add_fixedcatch_scorecard.R` and `R/configure_cmp_scorecard_oms.py` rebuild the OM-aware scorecard snapshots and interfaces. Input checksums and the R session are recorded in [input and software record](provenance.txt). This paper's [editable Quarto source](cmp-screening-working-paper.qmd) and companion data permit revision without editing the displayed figures manually.

`R/build_cmp_nearterm.R` rebuilds both sets of 2026–2035 plots and validates the reference Kobe points against the established reference figure.

Validation checks require 500 iterations for every OM/CMP/component summary, complete 15-by-26-year worm panels, common worm IDs, documented mass units, and direct numerical confirmation of every screening comparison. The HTML uses semantic headings and tables, numerical data alternatives, captions, and figure descriptions. Screen-reader testing and formal accessibility review remain outstanding. This paper analyses saved outputs using the existing model dynamics and reference-point conventions. Verification of the assessment and projection calculations remains separate work.


## Related working-group material

- [Full SC14 MSE working report](https://sprfmo.github.io/jmMSE26/evidence/sc14-mse-report.html): candidate definitions, conditioning, tuning and the broader results.
- [Technical paper and remaining work](https://sprfmo.github.io/jmMSE26/sc14/technical-paper.html): implementation and validation context.
- [Jack mackerel research wiki](https://sprfmo.github.io/SC14_JM/mse.html): management questions and links to the meeting papers.
