Transparency

Methodology & sources.

Every benchmark on EarnWell is built from a transparent blend of open datasets and anonymised user submissions. Here is exactly where the numbers come from and how they are computed.

Open data first

We seed every role × region cohort with credible open datasets so new users see useful numbers from day one.

Refreshed by users

Real submissions are weighted more heavily than reference data. The more the community shares, the more current the benchmark.

Privacy by design

We only show a cell when N ≥ 5 distinct submitters. Nothing that could re-identify a person is ever exposed.

Data sources

All reference rows are stored with source = 'reference' and are clearly distinguishable from live user submissions.

AI / ML & Data Salaries dataset
ai-jobs.net · 2023–2024
Source →
Coverage
~11,142 records across 60+ countries. Data, ML, AI and analytics roles.
Licence
CC0 1.0 (Public Domain)

Self-reported, anonymised. Used as reference baseline for tech IC and manager rows in AI/ML/Data families.

Stack Overflow Developer Survey
Stack Overflow · 2023
Source →
Coverage
~48,000 salary respondents globally. Software engineering, DevOps, SRE, mobile, web, data.
Licence
ODbL 1.0

Aggregated to role × level × country medians, then expanded into the same schema as user submissions.

Occupational Employment & Wage Statistics (OEWS)
U.S. Bureau of Labor Statistics · May 2024 release (latest)
Source →
Coverage
United States, all metropolitan statistical areas. Non-tech families (Sales, Marketing, Legal, Finance, HR, Support, Operations).
Licence
Public domain (U.S. government work)

Used to anchor US medians for non-tech roles. Extrapolated to other regions via cost-of-labour multipliers (see below).

Greek Startup Compensation Report 2025
Marathon Venture Capital · 2025
Source →
Coverage
Greece — tech, product, sales, ops across seed to Series B startups.
Licence
Publicly published report (educational use)

Used to strengthen the Athens / Thessaloniki cohorts and calibrate the SEE regional multiplier.

OECD Average Wages
OECD.Stat · 2024 update
Source →
Coverage
38 OECD member countries.
Licence
OECD Terms & Conditions (free reuse with attribution)

Used to derive cross-country cost-of-labour multipliers for extrapolating US non-tech medians globally.

Eurostat Structure of Earnings
European Commission — Eurostat · 2022 wave (latest quadrennial)
Source →
Coverage
EU-27 + EEA. Cross-country earnings by NACE sector and occupation.
Licence
Eurostat re-use policy (free with attribution)

Cross-check for European regional multipliers and sector deltas.

EarnWell user submissions
EarnWell community · Ongoing (rolling window)
Source →
Coverage
Growing — every role family, level and region users submit.
Licence
Contributor licence — anonymised aggregate use only

The core flywheel. As real submissions arrive they outweigh reference data in the recency-weighted blend.

How a benchmark is computed

  1. 1
    Cohort selection

    We select all submissions matching the requested role family, level, country and (optionally) metro and company size.

  2. 2
    Privacy floor

    If the cohort has fewer than 5 distinct submitters we show "Not enough data" rather than a number — no exceptions.

  3. 3
    Recency weighting

    Each row is weighted by how recently it was submitted. Rows older than 24 months are down-weighted; rows older than 36 months are dropped.

  4. 4
    Reference blending

    Where a cohort is thin, credible open-data reference rows fill the gap but are down-weighted vs. live user submissions.

  5. 5
    Percentiles

    We report p25 / p50 / p75 / p90 on total cash compensation (base + bonus). Equity is displayed separately when present.

  6. 6
    Currency

    Values are stored in the submitter's local currency and converted on the fly using a daily FX table. The user picks the display currency.

Validation & quality controls

Every submission — whether from a signed-in user or a bulk reference import — passes through the same server-side checks before it can influence a benchmark.

  1. 1
    Authenticated submission only

    User submissions require a signed-in session. The insert runs inside a server function that re-verifies the user's identity — the browser can't bypass it.

  2. 2
    Schema validation (Zod)

    Role family, level and region must be valid UUIDs pointing at rows that actually exist. The chosen level must belong to the chosen role family.

  3. 3
    Currency + FX sanity

    Currency must exist in our daily FX table. Base is converted to USD server-side and must fall between US$500 and US$5,000,000 — catches missing/extra zeros and wrong-currency mistakes.

  4. 4
    Range checks

    Local base must sit between 1,000 and 10,000,000 of the local currency. Equity is capped at 0–500% of base. Bonus is capped at 20,000,000 local. Years of experience 0–70.

  5. 5
    Duplicate + rate limit

    A user can submit at most 3 packages per 24 hours, and the exact same (role, level, region, base) tuple is rejected as a duplicate within 24h.

  6. 6
    Provenance flag

    Every row is stored with source = 'user' or source = 'reference'. Reference rows are down-weighted vs. live user submissions and are displaced as real data arrives.

  7. 7
    Privacy floor at read time

    Even if a cohort passes validation, the benchmark endpoint refuses to return percentiles until N ≥ 5 distinct submitters exist.

  8. 8
    Aggregates only

    Underlying rows never leave the server. Only p25 / p50 / p75 / p90 and histogram bucket counts are exposed to the client.

Employer contribution (bulk upload)

Paid employer plans require a one-time company-wide contribution before the full distribution unlocks. It's what keeps the dataset honest — everyone who reads also writes.

  1. 1
    Subscribe

    Pick a Starter or Team plan. Checkout completes and the workspace is provisioned immediately — but benchmark values and CSV export remain blurred until step 4.

  2. 2
    Prepare a file

    Download the CSV template from /employer/upload. Required columns: role, level, region, currency, base_amount. Optional: bonus_amount, equity_pct, years_experience, company_size, employment_type, notes. Names or slugs both work for role / level / region.

  3. 3
    Upload

    Drop in a CSV or .xlsx, or paste a Google Sheets share link. We parse client-side, show a preview grid, and flag missing columns before you commit.

  4. 4
    Server-side validation

    Each row runs through the same Zod schema, FX sanity, and range checks as an individual submission. Rows that fail are rejected with a reason; valid rows are inserted with source = 'employer_bulk' and linked to a batch record.

  5. 5
    De-identification

    We only accept role, level, region, comp components and coarse metadata. No names, emails, employee IDs, or free-text identifiers. Bulk rows are stored with user_id = NULL so they never appear in anyone's personal history.

  6. 6
    Access unlocks

    Once the first batch lands, the soft gate lifts across the employer dashboard — full percentiles, histograms, CSV export and seat management.

  7. 7
    Ongoing refreshes

    The upload page stays available. Later batches supersede older rows for the same role/level/region cohort, so your view tracks your live comp bands as they change.

Pay equity (Team plan)

Team-plan employers get a two-lens equity view over their uploaded roster. We deliberately don't require gender / ethnicity / age columns — the analysis works entirely from role, level, region and tenure so there's no protected-class data to store, leak or misuse.

  1. 1
    Market-relative position

    For every employee, we look up the EarnWell market cohort for their role / level / region and estimate their percentile by linear interpolation across P10 / P25 / P50 / P75 / P90. Cohorts under N=5 are shown as “—” rather than a misleading number.

  2. 2
    Below- and above-market flags

    Rows below the market P25 are flagged “below market”; rows above P90 are flagged “above market” so you can spot compression and outliers in the same pass.

  3. 3
    Internal cohort gap

    Within your own upload, we compute the median for each role / level / region cohort with at least 3 employees, then flag anyone more than 15% below or 25% above their own cohort median.

  4. 4
    Exportable PDF

    One-click PDF export includes the summary numbers, the full list of flagged rows, generation timestamp and methodology footnote. Suitable for board packs and internal HR review.

  5. 5
    No sensitive attributes stored

    We don't ask for and don't accept gender, ethnicity, age, name or employee ID. If your local regulator requires protected-class breakdowns (EU Pay Transparency Directive, UK gender pay gap, US EEO-1), use this report as an input alongside your own HRIS data — it is not itself a regulatory filing.

  6. 6
    Informational only

    Percentile estimates are directional. Not an offer, guarantee, or legal advice on pay equity compliance.

What this data is — and isn't

Reference rows are seeded from credible open datasets, not live market quotes from named employers.
Extrapolation to non-tech families outside the US uses OECD / Eurostat cost-of-labour multipliers; treat those cohorts as directional until user submissions accumulate.
Live accuracy improves with every submission — the reference layer is intentionally displaced as real data arrives.
EarnWell benchmarks are informational. They are not an offer, guarantee, or legal advice on pay equity.

Make the next benchmark more accurate.

Submit your anonymised compensation in under 90 seconds. You'll unlock the full dashboard and help the community.