CliMA / CliMA/CalibrateEmulateSample.jl

Advice/protection against oddities in training point sets

Open
#268 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Julia
Stars
90
Forks
16
PR merge metrics
No merged PRs in 30d

Description

Arising in PR #265 for example,

We find that sometimes emulator training is problematic for a fixed data set, and a small modification leads to massive improvements. More robust handling of the training dataest by e.g. providing more of a Cross validation procedure, or better construction of train/validation splits in the provided points may lead to more robust trainings.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing PR #265 and the emulator training path that consumes the provided training point sets. Investigate how train/validation splits are currently constructed and whether cross-validation is already supported. Done means defining and implementing a robust handling approach for fixed data sets, with evidence that small changes do not cause large unexplained training differences.

Written by the indexing model from the issue text.

Assessment

Tech stack
julia
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.