Potsdam Institute for Climate Impact Research (PIK)
Interdisciplinary Center for Dynamics of Complex Systems (University of Potsdam)
Cardiovascular Physics Group (Humboldt-Universität zu Berlin)
PIK Logo
TOCSY - Toolboxes for Complex Systems
PIK/ Antique/ Blue
Home Home
 ACE
 Adaptive Filtering
 Approx. RQA
 CoinCalc
 Commandline RPs
 COPRA
 COPRA2
 Coupling Analysis
 CRP Toolbox
 DSProlog
 Coupling Direction
 IOTA
 K2
 Makeinstall
 NEST
 PECUZAL
 PETROPY
 pyUnicorn
 RECFLOW
 RECGRAM
 RECLAC
 RP
 rqaci
 RSA
 System Identification
 TIGRAMITE
 SOWAS

COPRA2 – Constructing Proxy Records From Age Models

Release

(Standalone app, Python-cli, Python package)


General Notes

COPRA2 is a depth-age modeling program that creates chronologies with uncertainties and can transform age uncertainties to proxy uncertainties. COPRA2 is a modern Python port of the original MATLAB/Octave implementation of COPRA. It is backward compatible in some respects and comes with a number of differences and important improvements:

Backward compatibility

  • The results export includes a MATLAB-readable script (.m file), allowing results to be reused in the original MATLAB version of COPRA.

Differences

(full an detailed list at the end of the page)
  • NaN values in the input generally raise an error during data import; dropping them now requires explicit user confirmation.
  • The session export is now a JSON file. A MATLAB-compatible session export (the log file) can still be produced, but it no longer embeds the large data matrices; these are stored separately and loaded via a file-load command inserted into the exported .m file.
  • Interactive treatment of reversals and hiatuses has been improved: error bars can be modified individually, e.g., via a clickable multiplier or by entering an arbitrary value in an input field, or the hiatus line can be dragged to the desired position.

New features and improvements

(full an detailed list at the end of the page)
  • Reproducibility has been substantially improved: Monte Carlo sampling is now fully reproducible, since the random seed is stored in the session export file.
  • A computed age model can be reused for any other proxy record by loading a different proxy file into an open session.
  • Proxy errors can be included as an additional source of uncertainty.
  • The sample size of dating samples can be included as an additional source of uncertainty.
  • The algorithm now allows asymmetric age errors.
  • Improved robust hiatus detection.
  • Growth variability between the dating points removes the artificially small uncertainty between datings.
  • New tool for suggestions for additional dating points to improve age model or proxy uncertainty, based on simulation.
  • A preferences dialog in the GUI lets users set individual defaults for future sessions (e.g., confidence levels, interpolation method, mean vs. median), which are then applied automatically.
  • History of last opened sessions allows the user to quickly resume previous work.
  • In addition to the desktop app, COPRA is available as a command-line application and as a Python package, allowing easy integration of COPRA into Python scripts.

Introduction

COPRA transforms dating (age–depth) uncertainties into proxy uncertainties. Given a set of dated depths with age errors and a proxy record measured along the same depth axis, it builds a Monte Carlo ensemble of age–depth models, interpolates the proxy onto each realisation, and reports the proxy time series with confidence bands. Every run records its random seed, so results are exactly reproducible.

Compared with the original MATLAB toolbox, COPRA2 adds new uncertainty sources, checks and tools; see Differences from the MATLAB COPRA.

The software offers two front ends over one shared computational core: a desktop application (copra-gui) and a command-line tool (copra-cli).

Installation

From PyPI (with pip)

The distribution name is copra2; the import package and the console commands keep the copra name:

pip install copra2          # command-line tool + Python API (no GUI)
pip install "copra2[gui]"   # add the PySide6/Qt6 desktop application

PySide6 is an optional dependency. The command-line tool (copra-cli) and the copra.core API work without it, so old systems where PySide6 cannot be installed can still use COPRA; add [gui] to enable the desktop app.

From source (developers)

cd copra_py
python -m venv .venv
source .venv/bin/activate      # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
Use the virtual-environment interpreter (.venv/bin/python). A system Python with a mismatched NumPy build will fail to import the numerical core.

Standalone application (end users)

A double-clickable application is built with PyInstaller and needs no Python installation:

pip install pyinstaller
pyinstaller packaging/copra.spec     # run from copra_py/
PlatformResultLaunch
macOSdist/COPRA.appDouble-click. On first launch, right-click → Open to bypass Gatekeeper (unsigned build).
Windowsdist/COPRA/COPRA.exeDouble-click.
Linuxdist/COPRA/COPRARun the binary.

PyInstaller does not cross-compile; build on the target platform.

Launching the app

From a source checkout:

copra-gui                  # entry point
python -m copra.gui        # equivalent

Or double-click the standalone application built above.

Input data formats

Input files are plain text (.txt / .csv), one row per sample. Depths must be strictly increasing and unique — no duplicate depths, no NaN values.

  • Separators: spaces (also aligned, fixed-width columns), tabs, commas or semicolons, also mixed.
  • Decimal comma: numbers written as 45,2 (e.g. a German Excel export) are recognised when the columns are separated by semicolons, tabs or spaces; the log then says (decimal comma). If the comma is the only separator (120,980,45) it is always a separator.
  • Title and column-name lines at the top of the file (any line that does not start with numbers) are skipped, so files exported with a header row can be used as they are. Lines starting with # or % are comments and may appear anywhere. Text after the data ends the table.
  • COPRA refuses a file (with the line number) rather than guess when a text line sits between data rows, or when rows have different numbers of columns; fill missing values with NaN.
FileColumns (in order)Notes
Dating table (required) depth, [depth error,] age, age error [, upper age error] One row per dated level; age error is the 1σ uncertainty. An optional upper age error makes the error asymmetric (then the first age error is the lower one). An optional depth error in column 2 is the ± distance from the centre of the dating sample (e.g. 2 mm for a 4 mm wide sample). See Dating errors.
Proxy record (required) depth, [sample width,] proxy value [, proxy value error] Measured along the same depth axis as the dating table. Optional: the proxy value error (1σ) and the sample width (± half width of each sample). See Proxy errors.
Layer count (optional) depth, age, depth error Layer-count age is relative to the first counted layer (first layer has age 0).

Example dating table (depth in mm, age in yr, error in yr):

0     0      0
120   980    45
250   2010   60
...

With asymmetric errors (columns: depth, age, lower error, upper error):

1      50     12.5   20.5
112    207.9  20.4   19.4
...

Uncertainties

COPRA distinguishes two groups of uncertainty. Dating errors belong to the dating table and act when the Monte Carlo ensemble of age–depth models is built. Proxy errors belong to the proxy record and act afterwards, when the proxy is placed on the age axis of every realisation. All of them are optional except the age error.

ErrorGiven inMeaningDrawn as Acts on
Age errordating table (required)1σ of the dated age; optionally lower/uppernormal (split normal if asymmetric)Monte Carlo age model
Depth errordating table, column 2± distance from the centre of the dating sampleuniformMonte Carlo age model
Proxy value errorproxy file1σ of the measured proxy valuenormalproxy band
Sample widthproxy file or Sample width field ± half width of the depth interval a proxy sample integrates overuniformage of each proxy sample

All widths use the same ± convention: the distance from the centre of a sample to its edge, i.e. half the drill diameter or half the milled interval. All error values are in the units of their column (age errors in the age unit, depth errors and widths in the depth unit).

Dating errors

Age error. The uncertainty of each dated age, given as 1σ (one standard deviation; halve 2σ values first). In every realisation each age is drawn from a normal distribution around the dated age. With an additional upper age error the error is asymmetric: the age is drawn from a two-piece (split) normal, whose lower half has the lower error and whose upper half the upper error. The age errors are the only errors the treatment can change (error factors widen or narrow them, e.g. to make an age reversal tractable). Asymmetric errors have no MATLAB equivalent.

Depth error. A dating sample is not taken at a single depth but drilled or milled over a finite width, so its true position is only known to lie somewhere within the sampled interval. The depth error e is the ± distance from the centre of the dating sample to its edge: a sample 4 mm in diameter has e = 2 mm. It is given as the optional column 2 of the dating table (see column recognition).

How COPRA uses it: in every realisation the depth of each dating sample is drawn uniformly within depth ± e (every position within the sample is equally likely), together with its age drawn from the age error. The age–depth model of that realisation is then interpolated through these perturbed depths. A draw in which two dating samples swap their stratigraphic order is rejected and redrawn, just like a realisation that is not monotone in age. The depth error therefore widens the age uncertainty of the model, most where neighbouring dating samples lie close together compared with their size. Without a depth error (or with all values 0) the dating depths stay fixed, exactly as in the original COPRA.

The depth error is a property of the samples and is not changed by the treatment. It is shown as vertical error bars on the Age–depth plot (before and after the run); where a bar is too short to be seen at the plot scale, its value is added to the point label, e.g. 3 (±2). The treatment table shows the depth as depth ±e, and the Age–depth table tab has a column for it.

Between the dating points (optional)

Each realisation draws the ages at the dating points and connects them smoothly (linear, pchip or spline). Because the drawn ages are independent, the connecting curves average their errors: without further assumptions the age uncertainty is smallest in the middle between two dating points (for two equal errors the interval shrinks to about 71 %), although the growth rate there is actually unknown. The option Between datings → growth variability adds this growth-rate uncertainty:

  • In every realisation the time between two dating points is shared out over the gap by random, positive portions (a Dirichlet bridge with the smooth curve as its mean, varied at 8 equally spaced points per gap). The realisation still passes exactly through the drawn ages, stays monotone, and its growth rate varies within the gap. The scatter is largest in the middle of a gap and grows with the gap size.
  • How much the growth rate varies is estimated from the dating table, under one explicit assumption: the growth rate varies within a gap as much as it does from one dating interval to the next. The latter is measured from the logarithmic growth rates between consecutive dating points: s is derived from the differences of neighbouring rates (mean-square successive difference), so that a slow trend — e.g. fast growth at the top and slow growth at the bottom — is not mistaken for irregularity. Pairs across a hiatus are ignored. The strength of the bridge then follows from s without any free number (the time fraction spent in either half of a gap gets the same variance as for two independent, log-normally varying growth rates). The log reports s, the number of intervals it is based on and its uncertainty. The measured spread also contains the dating errors, so the estimate errs on the side of more uncertainty.
  • Off by default (smooth interpolation, as in the original COPRA). At least three dating points per segment are needed.
  • Whether the data support the option is tested with the uncertainty check, which can also calibrate its strength. A calibrated strength other than the estimate (e.g. 1.5×) is shown next to the check box and stored in the session. On the command line, --interp-uncertainty F sets the factor directly (0 = off, 1 = estimate); values other than 1 that do not come from the check are only a sensitivity test (“how much does my result depend on this assumption?”).

The option affects every realisation, so it carries through to all results: plots, tables and exported ensembles.

Checking the model uncertainty (leave-one-out)

Question: is the age uncertainty of the model realistic — neither too narrow (over-confident) nor too wide? COPRA answers it with the data themselves, by cross-validation: Tools → Check model uncertainty (leave-one-out) (CLI: copra-cli check-uncertainty).

Procedure.

  1. Every interior dating point — one with a neighbour above and below in the same hiatus segment — is left out once. (End points cannot be tested: the model would have to extrapolate.) The current treatment (removed points, error factors) and hiatuses are used; the layer count is not included.
  2. The age model is rebuilt from the remaining points (400 realisations) and the age at the left-out depth is predicted. The prediction contains the model spread plus the left-out point's own dating error (and its depth error, if any): it is the range of ages this dating could have shown if the model were right.
  3. The measured age is compared with the prediction by two measures:
    • Coverage — how many measured ages lie inside their predicted 95 % interval. For a realistic model about 95 %.
    • Spread ratio — the typical size of the standardised deviations z = (measured − predicted median) / predicted spread, computed robustly as 1.4826 × median |z| (equal to the standard deviation for normal deviations), so that a single strongly deviating point does not decide the check. About 1 is realistic, > 1 means the model is over-confident (intervals too narrow), < 1 too cautious. It uses all points, not only “in or out”, and is therefore more informative than the coverage. It is given with its own uncertainty (±).
  4. This is repeated with the between-dating variability off and at 0.5, 1, 1.5, 2, 3, 4 and 6× the estimate (same random numbers for all, so the settings are compared fairly). A setting fits if its spread ratio exceeds 1 by no more than its own uncertainty (i.e. it is not over-confident within the precision of the check). The suggested setting is: off if off fits; otherwise 1× (the estimate) if it fits — the estimate is then confirmed by the data; otherwise the smallest setting that fits.

Reading the result.

OutcomeMeaningWhat to do
Suggested: offThe model uncertainty already covers the left-out datings.Leave Between datings off.
Suggested: on (1×)Without it the model is over-confident between the dating points; the variability estimated from the growth rates is confirmed by the data.Click Apply suggested setting and run again.
Suggested: on, another factor (e.g. 2×)The estimate is not enough; this is the smallest setting that fits. Apply it; mention the calibration when reporting the results.
Not reachableEven at 6× the spread ratio stays above 1: the misfit is not (only) a matter of growth between the dating points, but rather of dating errors that are too small or of outliers. Look at the flagged points (|z| > 2.5, marked ⚠): widen their error or remove them in the treatment, then check again.

Points far off their prediction (|z| > 2.5, marked ⚠) are always listed, whatever the suggestion: they may be dating problems, or places where the growth rate changes abruptly. They do not decide the check.

The result window shows the reading in words, a plot of z for every left-out point against depth (grey band |z| < 1, dashed lines the 95 % limits), the settings compared, and a table per dating point (measured age, predicted median and interval, z). The summary is also written to the log.

Limits. The numbers are only as good as the number of points that can be left out: the spread ratio is uncertain by about ±1.2/√(2n) (±31 % for 7 points, ±15 % for 29); with few points the check can often not tell the settings apart, and it then does not ask for extra variability. With fewer than about six points treat the result as a rough indication (the window says so). Leaving a point out doubles the gap around it, so the check is somewhat stricter than the real situation. At least three testable points are needed.

Example (CLI) with a stalagmite record of 31 U/Th dates whose growth rate varies strongly (test data set many_dates):

copra-cli check-uncertainty --dating many_dates.txt

Leave-one-out check: 29 dating points, pchip, M=400
 between datings   covered     spread ratio
             off   20 / 29      1.47 +/- 0.23  over-confident
            0.5x   27 / 29      0.93 +/- 0.14  fits
              1x   27 / 29      0.89 +/- 0.14  <- suggested
            1.5x   28 / 29      0.62 +/- 0.09  fits
              ...
Suggestion: --interp-uncertainty 1 (the estimate is confirmed by the data)

Without the between-dating variability only 20 of 29 left-out dates lie inside their predicted 95 % interval: the model is clearly over-confident. With the variability estimated from the growth rates the check is passed.

Proxy errors

By default no proxy error is assumed. Both proxy errors are applied when the proxy is resampled onto the age-certain axis, after the Monte Carlo age model has been built; they do not change the age model itself.

Proxy value error. The measurement uncertainty of each proxy value, given as an optional proxy-file column. It must be given as 1σ (one standard deviation); if your lab reports 2σ, halve the values first. In every realisation each proxy value is perturbed by a normal distribution of that width, which widens the proxy confidence band.

Sample width. Like a dating sample, a proxy sample is drilled or milled over a depth interval, so its true depth lies somewhere within it. In every realisation the depth of each proxy sample is drawn uniformly within its interval and the sample's age is read there from the age–depth model, which adds an age (distance) uncertainty to the proxy record. It is set with the Sample width field in the GUI (CLI: --sample-width), in the depth unit of the files:

  • None — point samples (default);
  • Fixed ± — drilled holes of equal size: enter half the drill diameter (±0.5 for a 1 mm drill);
  • Continuous (milled) — contiguous samples: each reaches halfway to its neighbours, derived from the sample spacing (nothing to enter; CLI --sample-width continuous);
  • From file — a sample-width column (±) in the proxy file, for samples of different size. It always takes precedence; the mode is then set automatically.

No interval extends across a hiatus. Neither proxy error is present in the original MATLAB COPRA.

Depth error and sample width describe the same physical fact — a sample has a width — but for different samples: the depth error for the dating samples (it shapes the age model), the sample width for the proxy samples (it shifts where each proxy value sits on that age model).

Column recognition

The input files have no header, so COPRA recognises the meaning of the columns from their number and content. The log shows how a file was read, e.g. Dating table read as columns: depth, depth error, age, age error, and the table tabs list the columns.

Dating table

ColumnsMeaning
3depth, age, age error
4depth, age, lower age error, upper age error  or  depth, depth error, age, age error
5depth, depth error, age, lower age error, upper age error

With four columns COPRA decides from the data whether column 2 is the age or a depth error: ages rise (or fall) more or less monotonically with depth, apart from a few reversals, whereas depth errors are roughly constant or scatter around a typical value. The column that is clearly more monotone is taken as the age. If that is not conclusive, columns 3 and 4 are compared: two age errors are of similar size, while an age is usually much larger than its error. A negative column 2 is always an age (depth errors cannot be negative).

Proxy record

ColumnsRead as
2depth, value
3depth, value, value error  or  depth, sample width, value
4depth, sample width, value, value error

Column 1 is the depth. The proxy value carries the signal and scatters most, whereas a value error or a sample width is nearly constant; negative numbers can only be a proxy value. The Proxy table tab shows the file with its column assignment and a drop-down per column (Proxy value, Proxy value error, Sample width, ignore) to correct it; the file is then re-read and the assignment is saved in the session. When the columns cannot be told apart with certainty, COPRA warns in the log and opens that tab so the assignment can be checked. On the CLI use --proxy-columns, e.g. --proxy-columns sample_width value.

Desktop application

The window is split into a control panel (left), plot tabs (right) and a log/status area (bottom). Hover any input field for a tooltip.

  1. Select inputs. Choose the dating table and proxy record (and optionally a layer-count file). As soon as both required files are set, the review plots are drawn automatically.
  2. Set parameters. Number of Monte Carlo realisations, interpolation method, confidence-interval widths, an optional random seed, and the sample/proxy names (see Parameters).
  3. Review and treat. Inspect flagged age reversals and hiatus candidates; apply treatment if needed (see below).
  4. Run. The Monte Carlo model runs on a background thread with a progress bar; press Cancel to abort.
  5. Export. Write the results, ensemble, session and log to a chosen folder (File → Export results, or the Export button).

Menus

  • File — Open session (⌘O) reads a JSON session (all inputs, parameters, treatment and the random seed), so a run can be restored or reproduced. An open project is closed first, so nothing of it (results, treatment, plots) mixes with the loaded session. Save session (⌘S) writes it back to the same file without asking; only the first time (no file yet) it asks for a file name, like Save session as… (⇧⌘S), which always asks. The current session file is shown in the window title. Unsaved changes: when you close the session (⌘W), open another one, or quit COPRA with changes that are not saved (edited inputs, parameters or treatment, or a new run result), COPRA asks whether to Save, Discard or Cancel. Save only proceeds once the file has really been written; a just-loaded or just-saved session counts as unchanged. ⌘W first closes an open dialog window (e.g. the suggestions) and only then the session. Open recent lists the last ten sessions opened, saved or exported (missing files are left out; Clear menu empties the list). Replace proxy record… (⌘R), Export results… (⌘E), Close (⌘W), Quit. File dialogs start in the folder used last — separately for sessions, data files and exports — and remember it between sessions. (On Windows/Linux use Ctrl instead of ⌘.)
  • Settings → Preferences — appearance mode (system / light / dark) and default run parameters (MC realisations, interpolation, confidence intervals, output folder) and the result estimate, median (default) or mean. They are stored between sessions; the parameters prefill the fields on startup and the theme is applied immediately.
    The result estimate is used consistently for every result: the central age–depth curve, the age-certain age axis onto which the proxy is resampled, the central proxy series, the table tabs and the exported results table. Changing it after a run switches all of them at once (the proxy is resampled again; the Monte Carlo run is not repeated). The command line uses --stat mean|median (default median). The original MATLAB COPRA always used the mean age axis.
  • Help — open this documentation in the browser, or the project website.
  • About — version, authors, citation and license (About COPRA), and the Qt version (About Qt).

Plot and table tabs

TabShows
Age–depthThe editable review: dating points coloured by reversal severity, plus hiatus lines. All treatment (reversals, errors, hiatuses) is done on this plot. After a run it shows the median (or mean) age–depth model, with the dating points still labelled by their input ids.
Treated age model → Age model ensembleRead-only result of the current treatment; not interactive. After a run the tab is renamed Age model ensemble and shows all Monte Carlo realisations.
ProxyThe raw, measured proxy record versus depth (the resampled median/mean proxy series after a run).
RealisationsThe proxy plotted against every age realisation (the full ensemble as a time series).
Age–depth tableThe input dating table; after a run the median (or mean) age and the age confidence limits at each proxy depth.
Proxy tableThe input proxy record; after a run the median (or mean) proxy and its confidence limits on the age-certain axis (median or mean age, matching the result estimate). Selected cells can be copied with Ctrl/Cmd+C, e.g. into a spreadsheet.

When a reversal or an automatic hiatus is detected, a warning line prefixed with ⚠️ appears in the log/status area at the bottom.

Reversals & hiatuses

An age reversal is a dating point whose age is out of order relative to its depth. COPRA classifies reversals as tractable (resolvable by widening the age error) or non-tractable (best removed). A non-tractable reversal must be treated before a run: with it almost no Monte Carlo realisation is monotone, so the run could not finish. If the Monte Carlo run still finds no monotone realisation at all, it stops early with an explanation (likely causes: an untreated reversal, dating points too close in age for their errors, or spline overshooting — then try pchip or linear). Run stays disabled (its tooltip and the log name the pair of points) until a point of each such pair is removed or their errors are widened enough; the command line refuses to run likewise. Treatment is declarative and stored in the session, so a run configured in the GUI remains reproducible.

  • Pick reversal — click a dating point to toggle its removal (struck-through when marked). Removed points stay visible as hollow red circles, not connected to the age model, in the Treated age model tab and, after a run, in the Age–depth tab.
  • Pick error increase / decrease — click a point to widen or narrow its age error on a 0.5 grid (1.5×, 2×, 2.5×, …). Widening a point can turn a non-tractable reversal tractable; its colour and error bar update immediately.
  • Pick hiatus — on the Age–depth plot, click empty space to add a hiatus line, click a line to remove it, or drag a line to move it. A hiatus must lie strictly between two dating points.

All picking is done on the editable Age–depth plot; the Treated age model plot is the read-only result and is not interactive.

Automatic hiatus detection. COPRA flags gaps with an anomalously low growth rate as hiatuses. When one is found, its depth is written into the Hiatus depths field and drawn as a line on the Age–depth plot (and shown on the treated model); when none is found the field stays empty and no line is shown. Editing or clearing the field overrides the automatic value (auto-fill then stops). Note that an age reversal also produces a low-growth-rate gap, so a reversal can raise a spurious hiatus — treating the reversal removes it automatically.

Extrapolated ages and ages in the future

Extrapolation. Proxy samples above the first or below the last dating point get their ages by extrapolating the age model; the further away, the less reliable. Whatever the interpolation method, the age model is continued linearly there, with the growth rate of the outermost dating interval of each realisation. (Continuing the pchip or spline end polynomial, as the original COPRA did, bends away over longer distances and typically reverses — then no realisation is monotone and the run cannot finish.) COPRA warns in the log (e.g. Proxy extends 16.2 above the first dating (depth 0.3–16.5): its ages there are extrapolated) already when the data are loaded, and shades these ranges (hatched grey) in the Age–depth and Treated age model plots and, after a run, in the proxy time series.

Younger than today. An age younger than the present — the year COPRA is run, e.g. −76 yr BP in 2026 — is chronologically impossible. After a run COPRA checks the central age model: proxy samples younger than today are reported in the log with their depth range and marked in the plots (violet band, dash-dotted line at today's age). COPRA does not change these ages. The usual cause is an undated top section that is extrapolated over a long distance; check the youngest datings, add a dating near the top, or — for a stalagmite still growing when sampled — enter the top as a dating with the collection date. The check works for age scales in years or thousands of years BP (and b2k); for other scales it is skipped.

Suggesting additional datings

Tools → Suggest additional datings… (CLI: copra-cli suggest-datings) answers the question: where would one more dating improve the age model most? COPRA does not simply point at the widest confidence interval — it is always widest where the model is extrapolated and grows with age anyway. Instead it simulates the benefit of a hypothetical dating:

  1. About 40 candidate depths are spread over the proxy record (or over an allowed depth range you set, e.g. where datable material exists).
  2. At each candidate a hypothetical dating is inserted. Its age is the median of the current age model there; its error (1σ) is like that of the neighbouring datings (interpolated), a fixed value, or a percentage of the age — your choice.
  3. The age model is rebuilt (300 realisations, same random seed for all candidates) and its uncertainty measured: the mean 1σ spread of the ages over the proxy depths (Improve: age model), or of the proxy on the time axis (Improve: proxy record — age errors matter most where the proxy changes quickly).
  4. The best candidate is taken as set and the search repeated for the next suggestion (up to three), so suggestions do not cluster. Improvements below 1 % are Monte Carlo noise and are not suggested.

The result shows each suggestion (S1–S3) with its depth, expected age, assumed error, the reduction of the uncertainty and the reason (e.g. large gap between datings 11 and 12, extrapolated range below the last dating, steep proxy change nearby). A plot shows the suggestions on the age model and a benefit profile: the reduction a dating would bring at every depth — often more useful for choosing samples than three points, because it shows whether a whole range is worth sampling. Rule-based hints are added: a replicate dating between two points that form an age reversal, datings close to a hiatus, segments with only two datings.

The plot shows the age model of the last run (same statistic and confidence level as in the main window) when that run used the current data and settings; otherwise the model simulated for the suggestions (300 realisations, median and 95 %), as the legend says. Notes: the Between datings setting of the run is used (with it, large gaps count as uncertain, as they should); its strength is estimated once from the real datings. With pchip or spline a hypothetical dating very close to an existing one can even increase the uncertainty (it makes the curve locally more wiggly); such places are not suggested. The suggestions are hypotheses: COPRA does not know where datable material exists.

Reusing an age model for another proxy

A single Monte Carlo age model can be applied to several proxy records measured on the same core without re-running the (expensive) simulation. After a run, load a different proxy file in the proxy picker: the Run button changes to Transfer. Clicking it (or using File → Replace proxy record) evaluates the computed age–depth realisations at the new proxy's depths (linear interpolation per realisation) and updates the Proxy and Realisations tabs immediately; the age model itself is unchanged.

The new proxy should cover the same depth range; depths outside the age model's range are dropped. To keep the original results safe, COPRA renames the sample (appending the new proxy's file name) and requires a new sample name when you export, so nothing is overwritten. Start a completely fresh project with File → Close.

Command line

The copra-cli command drives the same pipeline without a GUI.

# Guided, interactive session (prompts for files, parameters, treatment)
copra-cli interactive          # or just: copra-cli

# Inspect a dataset: reversals + hiatus candidates, no simulation
copra-cli check --dating DATING.txt --proxy PROXY.txt

# Run the age model and write outputs (+ optional plots)
copra-cli run --dating DATING.txt --proxy PROXY.txt \
    --M 2000 --interp pchip --seed 42 --output-dir output --plots

# Reversal treatment: remove point 3, widen point 6's error three times
copra-cli run --dating DATING.txt --proxy PROXY.txt \
    --remove-points 3 --increase-error 6 6 6

# Layer counting and an explicit hiatus depth
copra-cli run --dating DATING.txt --proxy PROXY.txt \
    --layercount LAYERS.txt --hiatus 389.5

# Reproduce a previous run exactly from its saved session
copra-cli reproduce output/d_<sample>_<date>.json

Options may also be supplied through a TOML file via --config run.toml (a [copra] table with the same keys); explicit command-line options override the file.

Output files

Export first asks which result files to write; any combination can be selected, and the choice is remembered for the next export. Every result file is a .csv file with one header line. The field separator is chosen in the same dialog: comma (default), semicolon or tab; the last choice is kept between sessions (CLI: --delimiter comma|semicolon|tab). Numbers always use a decimal point, so for a spreadsheet set to a comma decimal separator (e.g. German locale) semicolon is the practical choice. Names carry the sample name and the export date in ISO YYYY-MM-DD form. The JSON session and the run log are always written.

Option / fileContents
Age model ensemble
<sample>_agemodel_ensemble_<date>.csv
Column 1 depth, column 2 the measured (original) proxy value, then one age column per Monte Carlo realisation.
Proxy ensemble
<sample>_proxy_ensemble_dt<ΔT>_<date>.csv
The proxy of every realisation on an equidistant age axis: column 1 age, then one proxy column per realisation. See below.
Proxy average
<sample>_proxy_<interp>_<date>.csv
Age, proxy (median or mean), age confidence limits, proxy confidence limits, depth: the classic COPRA results table, one row per proxy sample.
Growth rate
<sample>_growthrate_<date>.csv
Depth, age (median or mean), growth rate (median or mean) and its confidence limits. See below.
d_<sample>_<date>.json Canonical session: inputs, parameters and the random seed.
log_COPRA_<date>.m Legacy MATLAB-readable run log.

Proxy ensemble (equidistant age axis)

The proxy samples are irregularly spaced in age, and differently so in every realisation. The proxy ensemble puts all realisations on one common, equidistant age axis, e.g. for time-series methods that need even spacing:

  • The sampling time ΔT is entered in the export dialog (suggested: a round value near the typical sample spacing).
  • The axis runs from the minimum to the maximum age of all age models, rounded down / up to a multiple of ΔT.
  • For each realisation the measured proxy values are interpolated from that realisation's ages onto the axis, with the interpolation method of the run (linear, pchip or spline), separately for each hiatus-free segment. Ages outside a realisation's range, or inside a hiatus gap, are left empty (nan).

Growth rate

The growth (deposition) rate is the slope of the age–depth curve, depth per time (e.g. mm/yr). It is computed for every realisation at every proxy sample from the neighbouring samples:

growth rate = (di+1 − di−1) / (ti+1 − ti−1)

with depth d and age t. The first and last sample of a segment use the one-sided difference (d2 − d1) / (t2 − t1). Nothing is differenced across a hiatus; where no time elapses the rate is undefined (nan). The file gives the median (or mean) of the realisations' rates and their quantiles for the age confidence level.

Reproducibility

The JSON session is the authoritative record of a run. It stores the input file paths, all parameters and the resolved random seed. Feeding it back reruns the model to a bit-identical ensemble:

copra-cli reproduce output/d_sample_2026-09-13.json

If no seed is given, COPRA generates one and records it, so even “random” runs remain reproducible after the fact.

Parameters

ParameterMeaningDefault
MC realizations (--M) Number of Monte Carlo age models. More gives smoother statistics but is slower.2000
Interpolation (--interp) linear, pchip (monotone) or spline, used between the first and the last dating point. pchip avoids overshoot. Outside the dated range the age model is always continued linearly with the slope of the outermost interval (see Extrapolated ages). pchip
Proxy CI (%)Confidence-interval width for the proxy band (e.g. 95 → 2.5–97.5% quantiles).95
Age CI (%)Confidence-interval width for the age band. 95
Between datings (--interp-uncertainty)Growth-rate variability between the dating points: off, or on (= estimated from the data; a factor from the uncertainty check). See Between the dating points.off
Seed (--seed)Random seed; blank/omitted draws a fresh one and records it. In the GUI the seed used is shown in the field after each run and reused by later runs; it is cleared when a new dating or proxy file is loaded (or clear it yourself for a fresh seed). random
Sample nameUsed in output file names. from dating file
Proxy nameAxis label on the plots. δ18O

Python API

The scientific core is a plain NumPy/SciPy package with no Qt or matplotlib dependency and can be scripted directly:

from copra.core.model import CopraSession
from copra.core.logging_ import run_session, save_results

session = CopraSession(
    dating_path="DATING.txt",
    proxy_path="PROXY.txt",
    M=2000,
    interp_method="pchip",
    samplename="mycore",
    seed=42,
)
run_session(session)                       # populates the ensemble + statistics
save_results(session, "output")            # writes the four output files

Package layout: copra.core (computation), copra.plotting (matplotlib figures), copra.gui (Qt6 application), copra.cli (command-line tool).

Differences from the MATLAB COPRA

COPRA2 is a complete reimplementation of the MATLAB/Octave toolbox. The scientific core — Monte Carlo age–depth modelling with monotonicity by rejection, resampling of the proxy onto an age-certain axis — is the same, and on the original test data the results agree statistically (Monte Carlo runs are never identical bit for bit, since the random numbers differ). Beyond that, COPRA2 adds features and fixes weaknesses of the original. The technical details, with the MATLAB files concerned, are recorded in DIFFERENCES.md in the source code.

New scientific features

  • Asymmetric age errors (separate lower and upper error), sampled with a split normal distribution. Input data
  • Depth error of the dating samples (± half the sample width): the depth of each dating is drawn uniformly within it in every realisation. Dating errors
  • Proxy errors: a 1σ proxy value error, and the sample width of the proxy samples (none, fixed ±, continuous/milled, or per sample from the file). Proxy errors
  • Growth variability between the dating points (optional): removes the artificially small uncertainty between datings of the original, with its strength estimated from the data. Between the dating points
  • Uncertainty check (leave-one-out): tests whether the model uncertainty is realistic and whether the growth variability is needed. Uncertainty check
  • Suggestions for additional datings, with a benefit profile over depth. Suggest datings
  • Growth rate with confidence limits, and a proxy ensemble on an equidistant age axis, as outputs. Output files
  • Reusing an age model for another proxy record of the same sample, without a new Monte Carlo run. Reuse age model

Changed behaviour (results can differ from MATLAB)

  • One statistic for all results: age axis, age model, proxy series, tables and export all use the same median (default) or mean. MATLAB always used the mean age axis and mixed it with a median proxy.
  • Linear extrapolation outside the dated range, for every interpolation method. MATLAB continued the pchip/spline end polynomial, which bends away and often made a run impossible. Extrapolated ages
  • No age reversal across a hiatus: the segments above and below a hiatus are drawn together, so no realisation is older above the hiatus than below it. MATLAB drew them independently.
  • Hiatus detection ignores the (negative) growth rates of age reversals, which used to cause spurious hiatus candidates.

Checks and safety

  • Untreated non-tractable reversals block the run (the MATLAB code would loop for a very long time); a run that cannot find any monotone realisation stops early with an explanation.
  • Ages younger than today are reported and marked, and extrapolated depth ranges are shaded. Plausibility
  • Stricter, more tolerant file reading: decimal commas, comment lines anywhere, recognition of optional columns (depth error, sample width, value error) with a manual column assignment when unclear; no silent loss of data (incomplete rows, text between data, rows of different length are reported with their line numbers). Input data, Column recognition
  • Missing values (NaN) are rejected by default instead of being dropped silently; the program asks whether to remove the rows.

Usability

  • A new desktop application (Qt6): all treatment (reversals, error factors, hiatuses) is done directly on the Age–depth plot; removed datings stay visible; table tabs for the dating table, the proxy record and the results; runs in the background with progress bar and cancel.
  • Sessions: open, save / save as, open recent, remembered folders, keyboard shortcuts; preferences are kept between sessions; the program starts quickly (heavy modules are loaded when first needed).
  • A command-line tool for scripted and batch use, with the same features, and a Python API.

Output and reproducibility

  • Every run records its random seed; a saved session reproduces a run exactly. The session is a readable JSON file (instead of a MATLAB .mat file); a MATLAB-style .m log is still written.
  • Export of four selectable result files, all .csv with a header line and a selectable separator; ISO dates (YYYY-MM-DD) in the file names. Output files

Download

Last version (v2.0.0-rc.8):

COPRA — macOS (Intel x86_64) · macOS 15.0 or newer
COPRA — macOS (arm64) · macOS 15.0 or newer
COPRA — Windows (x64) · Windows 10 or newer
COPRA — Linux (x86_64) · glibc 2.39 or newer

Source: https://gitlab.com/tocsy/copra
Commandline use: pip install copra2


References

  • Breitenbach, S. F. M., Rehfeld, K., Goswami, B., Baldini, J. U. L., Ridley, H. E., Kennett, D., Prufer, K., Aquino, V. V., Asmerom, Y., Polyak, V. J., Cheng, H., Kurths, J., Marwan, N.: COnstructing Proxy Records from Age models (COPRA), Climate of the Past, 8, 2012, 1765-1779
    doi:10.5194/cp-8-1765-2012.


Authors

Norbert Marwan, Sebastian Breitenbach, Claude Code Sonnet 5
Acknowledgements (original version): Kira Rehfeld, Bedartha Goswami, Daniel Juncu.



Disclaimer on the use of generative AI

The original MATLAB implementation was developed without generative-AI assistance. The port to Python and the Qt6 implementation were developed with the assistance of Anthropic's Claude Sonnet 5. All AI-assisted output was checked against the original MATLAB implementation to verify correct reproduction of the underlying algorithm.


License

This project is licensed under the GNU General Public License v3.0.


© 2004-2026 SOME RIGHTS RESERVED
University of Potsdam, Interdisciplinary Center for Dynamics of Complex Systems, Germany
Potsdam Institute for Climate Impact Research, Complexity Science, Germany

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 2.0 Germany License.
Imprint, Data policy, Disclaimer, Accessibility statement

Please respect the copyrights! The content is protected by the Creative Commons License. If you use the provided programmes, text or figures, you have to refer to the given publications and this web site (tocsy.pik-potsdam.de) as well.

@MEMBER OF PROJECT HONEY POT
Spam Harvester Protection Network
provided by Unspam