openstudy.id

How to contribute

Routed by what you already have, not by what you do for a living — a quantitative economist is a researcher and a data scientist both, and picking one would be a fiction. Each path below carries its honest capacity, because one person cannot service five doors at once.

Ways infind your row
What you already haveThe pathStatus
A published paper, and the data behind it An instrument built on it open
A subject you know, and no instrument Interpret one of the 10 waiting open
A question and a data source, no analysis yet A new study, gated before anything is built open
Skills and time, no particular subject Replicate or break an existing instrument case by case
An instrument you built, and no domain author Bring it, with a named target case by case

Case by case is not a soft no — both are real and both have happened. It means capacity is the constraint, so ask before investing time rather than after.

Path onea published paper

The research is already refereed. Now make it usable.

A paper's analysis is its most reusable part and its least accessible: the figures are flattened into a PDF, and the question a reader wants to ask — what happens if I change this assumption? — cannot be asked. That is what gets built here, on top of work that has already passed review.

This is the cleanest path on the site, because the objections that apply to the others do not apply to it:

Why it is clean

  • No preprint problem. Your paper is published. Nothing here affects a submission you have already made
  • No peer-review caveat. The research is refereed; the instrument is an addition to it, not a substitute
  • Authorship is settled. No negotiation — the instrument is a separate artifact that cites your paper
  • It is fast. The hard interpretive work is done, so there is no waiting on anybody

Two things that stop it dead

  • The data. If you cannot share the data behind the paper, there is no instrument. This is the first question asked, and it ends more of these conversations than anything else
  • Publisher copyright. Findings and data are yours to re-express, but the journal usually holds the typeset figures and text. Everything is rebuilt from the data; the published figures are never reproduced

It is not a figure service. What gets added is not a prettier chart: the pipeline is re-runnable from source, the models are scored against baselines the paper probably did not report, and the gates are published including any that fail. If re-running your analysis honestly weakens a result, the instrument says so — which is the same deal everyone else here gets.

Path twoa subject you know

10 instruments are built and waiting for a conclusion

No data to find, no code to write, nothing to build. The measurement is finished on all of them; what is missing is what the result means in its own field, and each instrument states its own specific gap rather than a generic invitation.

Write that interpretation and you are first author. You are also accountable for it, which is the whole point — the reason the pages currently stop short of a conclusion is that nobody involved is qualified to draw one.

See the 10 openings, each with what it found and what it is missing.

Path threea question and data

Starting from a question, with gates before anything is built

The most valuable path and the easiest to abandon halfway, so it is deliberately gated. Nothing gets built until the first three steps are done in writing.

What happens, in order

  1. You say what the question is and point at the data. A conversation, not a form.
  2. I scope it — what can be built from open data, what cannot, and where the honest limits are. If it cannot be built well, you are told then, not later.
  3. Thresholds are agreed before anything runs. What would count as the model working, and what would count as it failing. Written down first.
  4. The instrument gets built. You see it while it is being built, not at the end.
  5. You write the interpretation. What the result means, what surprised you, what it does not settle.
  6. It publishes with your name as lead author, and a DOI if you want one.

Step 3 is the one that matters. Thresholds are never moved to fit an outcome. If the instrument does not support your finding, the page says so, with the numbers. That is not a threat — it is the reason a reader should believe the cases where it does.

Open data only, aggregate only, no human subjects. Not a judgement about your work — a small operation cannot carry ethics review or personal data, and the limits below say so plainly.

Case by caseskills, no subject

Replicate or break something that already exists

Per hour spent this is the most valuable contribution anyone can make here, and it is case by case only because reviewing it well takes my attention rather than because it is unwelcome. An instrument that a named engineer attacked and failed to break is worth more than three nobody checked.

Break the evaluation — a day

Is the metric right for the problem? Did a baseline see the test set? Does the cross-validation actually prevent leakage? If you succeed, that gets published — with Validation credit.

Beat the baseline — a week

Every model here is scored against a naive rival that already contains the easy knowledge. Build a better rival. If yours wins, the case is weakened and that is published; if it loses, the case is stronger. Either outcome is worth having.

Three specific openings right now:

  • transit-equity has two failing hard gates — timetable sanity and network integrity. Are they fatal or cosmetic?
  • rice-security's R² goes from −11.09 to 0.82 through an OLS calibration. Is that legitimate, or is it fitting the answer?
  • fire-haze scores a 3.8% event class with ROC-AUC. Is precision-recall the honest metric instead?
Case by caseyour own instrument

Bringing an instrument you built yourself

This works — every instrument here was built before it had a domain author. But the constraint is worth stating plainly: there are already 10 instruments without one. The shortage is domain authors, not instruments, and another unpartnered build makes that worse rather than better.

So it is welcome with a named target — a person or an institution you have already approached about interpreting it. Bring that and the conversation is short. Without it, the honest answer is that it would join a queue that is not moving.

It also has to meet the published standard, which includes the visualization criteria — the part most often skipped.

CreditCRediT, not authorship

Naming what each person did

The Contributor Roles Taxonomy replaces "author or not" with fourteen named roles. It is a published standard, used by many journals, and it is the honest way to describe this arrangement — and the reason paths are routed by contribution rather than by job title.

RoleDomain authorMeasurement
Conceptualizationyes
Investigationyes
Writing – original draftyes
Resources — domain data, access yes
Methodologysharedyes
Formal analysisyes
Validationyes
Softwareyes
Data curationyes
Visualizationyes

Six roles including Methodology and Formal analysis is, under the criteria most journals apply, a methods co-authorship — designing the analysis and evaluating the models is not a technical service. What it does not claim is the finding: you lead, you own the interpretation, and you are accountable for what the result means.

Author order. Whoever writes the interpretation is first author. Whoever built the measurement is second, as methods co-author.

Stated here so it is never a negotiation later. If a case has several contributors, order follows contribution and is agreed before publication — never after, and never by seniority.

Where the site is merely run rather than contributed to, that is an editor credit and not authorship — the same rule applied to the person who runs it.

AI assistance, disclosed

Pipeline and model implementation are produced with AI assistance, under my direction. It is stated on every instrument, in the methods, in a sentence — because a reader who works it out for themselves is entitled to feel misled, and a reader who is told is not.

Limitswhat is not accepted

What this site will not publish

  • Personal data. Aggregate only, no exceptions.
  • Redistributed source data. Pipelines fetch from the origin and cite the licence; the instrument publishes derived results, not somebody else's dataset.
  • Accusations against named companies or projects. Methods and aggregate distributions, yes. Naming an actor as culpable, no.
  • A finding without a way to check it. If the gates cannot be stated, the instrument cannot be built.

These are limits on the site, not judgements about your work. Some of them exist because a small operation cannot carry the risk; the rest are on the standards page with the reasoning.