L170open world research agent
discover, theorize, experiment, revise and report over a long horizon
languageenglishsamples50tagstransfer, capstone, open-worlddifficulty knobnoambiguity5discourse horizon5lexical novelty5reasoning depth5world complexity5capabilitiesopen_ended_discovery, scientific_induction, metareasoning, self_modeling, architecture_adaptation
seed 000english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 000lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 001lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 002lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 003lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 004lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 005lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 006lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 007lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 008lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 009lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 010lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 011lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 012lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 013lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 014lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 015lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 016lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 017lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 018lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 019lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 020lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 021lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 022lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 023lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 024lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 025lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 026lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 027lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 028lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 029lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 030lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 031lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 032lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 033lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 034lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 035lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 036lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 037lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 038lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 039lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 040lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 041lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 042lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 043lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 044lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 045lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 046lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 047lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 048lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049english
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049spanish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049chinese
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049turkish
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049swahili
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049english_synonym
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049symbols
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049deu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049sh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049fin
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049cmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049jpn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049tur
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049arb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049kat
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049mri
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049vie
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049tha
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049tam
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049swh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049nav
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049kal
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049tpi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049unm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049inh
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049kor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049eus
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049chr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049quz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ady
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049mww
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049wyi
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049cic
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ain
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ckt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049yua
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049twf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049lkt
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049gug
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ayr
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049nmn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049niv
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049lut
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049yag
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049luo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ppl
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049fra
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049epo
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049hun
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049yue
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ryu
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049kaz
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049heb
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049xmf
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ind
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049khm
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049lao
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049tel
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049yor
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049apw
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049ike
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049srn
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049chy
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.
seed 049lez
LessonNotImplemented: open_world_research_agent: Not reducible to a one-step exactly-graded episode, and not faked. The lesson's content is a *loop*: discover an ontology, acquire the language that describes it, propose competing theories, DESIGN an experiment, run it, revise under criticism, and report with provenance. Everything that distinguishes it from lessons 150-169 lives in the parts a single question cannot contain: (1) the agent must choose interventions, so the environment has to accept experiment programs as actions and return outcomes, i.e. a multi-step interactive environment with a hidden generative world model rather than a generate()->answer pair; (2) the score is a trajectory functional — sample efficiency, whether each claim is supported by an experiment the agent itself ran, calibration of stated confidence, and whether a refuted theory is actually abandoned — none of which is a function of one answer; (3) grading requires an adversarial critic and a provenance ledger, both of which are additional environments. A real implementation needs: a parameterized hidden world (latent laws + nuisance parameters) with an exact simulator; an action language for experiments, assertions and retractions; an evaluator that checks each asserted law against the true one and each claim against the agent's own evidence ledger; and a scalar built from (discovery completeness x provenance validity x calibration) over the whole trajectory. Until that harness exists, any single-step version would grade a different, easier lesson while wearing this one's name.