Predictive Modeling

Final Analytics: What Is Associated With Higher Pay, and Which Role Is the Best Opportunity?

Overview

The earlier pages describe the market. This page turns that description into decision support for a job seeker on the Data Analyst to Data Scientist pathway in NAICS 5182. It has two analyses:

  1. A salary regression (Part 1) that asks which factors are associated with higher disclosed pay.
  2. A role opportunity scorecard (Part 2) that ranks the four pathway roles on pay, how accessible they are, and how well they fit our team’s current skills.
  3. A salary estimator (Part 3) where you enter a posting’s details and the model estimates the salary and explains what drives it. It is trained on a broader file than Parts 1 and 2, as Part 3 explains.

Both use the MET Career Compass 2026 job postings (the course Jobs_2026 folder) described in Data Preparation: 765 pathway postings in NAICS 518, of which 303 disclose a salary and 570 include posting text. The build script is build_met_text_panel.py. Each part states what was modeled, why it matters, which variables were used, how it was evaluated, and what the result means in practical terms.

import re
import warnings

import numpy as np
import pandas as pd
import plotly.graph_objects as go
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.model_selection import KFold, RepeatedKFold, cross_val_predict, cross_validate
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder

warnings.filterwarnings("ignore")

PALETTE = ["#457b9d", "#2a9d8f", "#e9c46a", "#e76f51", "#6d597a", "#f4a261"]
GRAY = "#b8bec6"


def style(fig, title, subtitle, height=450, legend=False):
    fig.update_layout(
        title=dict(text=f"<b>{title}</b><br><span style='font-size:12px;color:gray'>{subtitle}</span>",
                   x=0, xanchor="left"),
        template="simple_white", height=height, showlegend=legend,
        font=dict(family="Segoe UI, Helvetica, Arial, sans-serif", size=13),
        margin=dict(l=10, r=20, t=95, b=50), hoverlabel=dict(bgcolor="white", font_size=13),
        legend=dict(orientation="h", x=0.5, xanchor="center", y=-0.15, yanchor="top"))
    fig.update_xaxes(gridcolor="#eceff3")
    fig.update_yaxes(gridcolor="#eceff3")
    return fig


def show(fig):
    fig.show(config={"displaylogo": False, "responsive": True, "modeBarButtonsToRemove": ["lasso2d", "select2d"]})


mp = pd.read_csv("data/processed/met_text_panel.csv")
mp["salary_mid"] = (mp["salary_min_annual"] + mp["salary_max_annual"]) / 2
SENIOR = r"senior|\bsr\b|staff|principal|\blead\b|manager|director|head of|architect"
mp["level"] = np.where(mp["title"].str.lower().str.contains(SENIOR), "Senior title", "Other title")
state = mp["state"].fillna("")
mp["state_group"] = np.where(state.isin(["California", "New York", "Virginia", "District of Columbia"]), state,
                             np.where(state == "", "State not given", "Other states"))
mp["remote"] = mp["remote_status"].fillna("Unknown").replace({"Unknown": "Not stated"})
print(f"{len(mp)} postings, {mp['salary_mid'].notna().sum()} with a salary, {int(mp['has_text'].sum())} with posting text")
765 postings, 303 with a salary, 570 with posting text

Part 1: What Is Associated With Higher Pay?

What was modeled. The annual salary of a posting, taken as the midpoint of its disclosed range.

Why it matters for a job seeker. The pay gap between roles and places is the main reason to move along the pathway. A model shows which factors are still associated with pay once the others are held fixed, for example whether a New York posting pays more than a California one for the same role.

Variables used. Four features, all read from the posting: the role (Data Analyst, Data / Analytics Engineer, Data Scientist, or ML Engineer, from the job title), whether the title signals a senior level (senior, staff, principal, lead, manager, director, head of, or architect), the state group (California, New York, Virginia, District of Columbia, other states, or no state given), and the work arrangement (remote, hybrid, onsite, or not stated). Experience is not in these files, and skills are available for only about half of the postings with a salary (158 of 303), so they are not used here.

How it was evaluated. We compared three models with repeated 5-fold cross-validation (5 repeats), so every score comes from postings the model had not seen: a baseline that always guesses the median salary, a ridge regression (a linear model that shrinks coefficients to avoid overfitting a small sample), and a random forest. The metrics are RMSE and MAE (typical error in dollars) and R² (the share of the variation in pay that the model explains; a value at or below zero means no better than guessing the average).

sal = mp.dropna(subset=["salary_mid"]).copy()
features = ["role", "level", "state_group", "remote"]
X, y = sal[features], sal["salary_mid"].values


def pipe(model):
    pre = ColumnTransformer([("cat", OneHotEncoder(handle_unknown="ignore"), features)])
    return Pipeline([("prep", pre), ("model", model)])


models = {
    "Baseline (median)": DummyRegressor(strategy="median"),
    "Ridge regression": Ridge(alpha=3.0),
    "Random forest": RandomForestRegressor(n_estimators=300, min_samples_leaf=3, random_state=0),
}
cv = RepeatedKFold(n_splits=5, n_repeats=5, random_state=0)
rows = []
for name, model in models.items():
    r = cross_validate(pipe(model), X, y, cv=cv,
                       scoring=("neg_root_mean_squared_error", "neg_mean_absolute_error", "r2"))
    rows.append({"Model": name,
                 "RMSE (USD)": -r["test_neg_root_mean_squared_error"].mean(),
                 "MAE (USD)": -r["test_neg_mean_absolute_error"].mean(),
                 "R2": r["test_r2"].mean()})
results = pd.DataFrame(rows).set_index("Model")
results.round({"RMSE (USD)": 0, "MAE (USD)": 0, "R2": 2})
RMSE (USD) MAE (USD) R2
Model
Baseline (median) 61849.0 48183.0 -0.04
Ridge regression 50254.0 37888.0 0.31
Random forest 49785.0 37960.0 0.32
fig = go.Figure()
for col, color in [("RMSE (USD)", PALETTE[0]), ("MAE (USD)", PALETTE[1])]:
    fig.add_trace(go.Bar(
        x=results.index, y=results[col], name=col, marker_color=color,
        text=[f"${v / 1000:,.1f}K" for v in results[col]], textposition="outside", cliponaxis=False,
        hovertemplate="<b>%{x}</b><br>" + col + ": $%{y:,.0f}<extra></extra>"))
fig.update_layout(barmode="group")
fig.update_yaxes(title="Typical prediction error (USD)", tickprefix="$", tickformat=",.0f",
                 range=[0, results["RMSE (USD)"].max() * 1.18])
style(fig, "How Well Each Model Predicts Salary",
      f"Cross-validated error on {len(sal)} salary-disclosed postings | lower is better", legend=True)
show(fig)

Interpretation: the ridge regression and the random forest are essentially tied, and both clearly beat the baseline. Ridge’s typical error (RMSE) is about $50.3K against $61.8K for the baseline, its mean absolute error is $37.9K against $48.2K, and its R² is 0.31 against -0.04 for the baseline. The random forest scores almost the same (RMSE $49.8K, R² 0.32), a difference too small to matter, so we use the simpler, more explainable ridge model, whose coefficients we can read and plot. In practical terms, role, seniority, state, and work arrangement together explain about 31% of the differences in pay between postings. The model is useful for comparing factors, but a typical estimate is still about $38K away from the real figure, so it should not be used to predict a specific offer.

final = pipe(Ridge(alpha=3.0)).fit(X, y)
names = final.named_steps["prep"].get_feature_names_out()
coef = pd.Series(final.named_steps["model"].coef_, index=names)

# 400 bootstrap refits give a 90% range for each factor
rng = np.random.default_rng(0)
boot = []
for _ in range(400):
    idx = rng.integers(0, len(sal), len(sal))
    m = pipe(Ridge(alpha=3.0)).fit(X.iloc[idx], y[idx])
    boot.append(pd.Series(m.named_steps["model"].coef_, index=m.named_steps["prep"].get_feature_names_out()))
boot = pd.DataFrame(boot).reindex(columns=coef.index).fillna(0)
eff = pd.DataFrame({"effect": coef, "low": boot.quantile(0.05), "high": boot.quantile(0.95)})
eff.index = [i.replace("cat__", "") for i in eff.index]

show_rows = {
    "role_ML Engineer": "Role: ML Engineer",
    "role_Data Scientist": "Role: Data Scientist",
    "role_Data / Analytics Engineer": "Role: Data / Analytics Engineer",
    "role_Data Analyst": "Role: Data Analyst",
    "state_group_New York": "State: New York",
    "state_group_California": "State: California",
    "state_group_Other states": "State: other states",
    "state_group_District of Columbia": "State: District of Columbia",
    "state_group_State not given": "State: not given",
    "state_group_Virginia": "State: Virginia",
    "level_Senior title": "Senior title",
    "remote_Remote": "Remote posting",
    "remote_Hybrid": "Hybrid posting",
    "remote_Onsite": "Onsite posting",
}
plot = eff.loc[list(show_rows)].rename(index=show_rows).sort_values("effect")
group_color = [PALETTE[3] if i.startswith("Role") else PALETTE[0] if i.startswith("State")
               else PALETTE[1] for i in plot.index]

fig = go.Figure(go.Scatter(
    x=plot["effect"], y=list(plot.index), mode="markers",
    marker=dict(size=11, color=group_color),
    error_x=dict(type="data", symmetric=False, array=plot["high"] - plot["effect"],
                 arrayminus=plot["effect"] - plot["low"], color="#6b7280", thickness=1.5, width=4),
    customdata=np.stack([plot["low"], plot["high"]], axis=1),
    hovertemplate="<b>%{y}</b><br>Effect: %{x:+,.0f} USD<br>90% range: %{customdata[0]:+,.0f} to %{customdata[1]:+,.0f}<extra></extra>"))
fig.add_vline(x=0, line_dash="dash", line_color="#6b7280")
fig.update_xaxes(title="Difference from an average posting (USD per year)", tickformat="+,.0f", zeroline=False)
style(fig, "What Is Associated With Higher or Lower Pay",
      "Ridge regression, other factors held fixed | dots are estimates, lines are 90% bootstrap ranges", height=560)
show(fig)

Interpretation: role is by far the strongest factor. Compared with an average posting (about $195K), an ML Engineer title is associated with about $31K more and a Data Scientist title with about $21K more, while a Data Analyst title is associated with about $51K less. A Data / Analytics Engineer title is about the same as average, because its range crosses zero. Location matters too: California is associated with about $25K more and New York with about $15K more, while Virginia is about $24K lower and the District of Columbia about $12K lower; the other state groups cross zero. A senior title adds about $14K relative to an average posting, or about $28K more than other titles. Work arrangement is mixed: hybrid postings are about $15K above average with a range that stays above zero, though they rest on only 25 salary postings; remote postings are about the same as average, because their range crosses zero; and onsite postings are about $27K below average, but only 10 postings are onsite, so that estimate is fragile. For a job seeker, the clearest signal is that moving from a Data Analyst title toward a Data Scientist or ML Engineer title is associated with a pay difference of roughly $72K to $82K, larger than any location choice (California against Virginia is about $49K).

oof = cross_val_predict(pipe(Ridge(alpha=3.0)), X, y, cv=KFold(5, shuffle=True, random_state=0))
role_colors = {"Data Analyst": PALETTE[0], "Data / Analytics Engineer": PALETTE[1],
               "Data Scientist": PALETTE[3], "ML Engineer": PALETTE[4]}

fig = go.Figure()
for role, color in role_colors.items():
    mask = (sal["role"] == role).values
    fig.add_trace(go.Scatter(
        x=oof[mask], y=y[mask], mode="markers", name=f"{role} (n={mask.sum()})",
        marker=dict(size=8, color=color, opacity=0.75),
        customdata=sal.loc[mask, "title"],
        hovertemplate="<b>%{customdata}</b><br>Predicted: $%{x:,.0f}<br>Actual: $%{y:,.0f}<extra></extra>"))
lims = [float(min(oof.min(), y.min())) - 10000, float(max(oof.max(), y.max())) + 10000]
fig.add_trace(go.Scatter(x=lims, y=lims, mode="lines", line=dict(dash="dash", color="#6b7280"),
                         name="Perfect prediction", hoverinfo="skip"))
fig.update_xaxes(title="Predicted salary (USD, from postings the model had not seen)", tickprefix="$", tickformat=",.0f")
fig.update_yaxes(title="Actual midpoint salary (USD)", tickprefix="$", tickformat=",.0f")
style(fig, "Predicted vs. Actual Salary",
      f"Cross-validated ridge predictions | correlation {np.corrcoef(oof, y)[0, 1]:.2f}", height=520, legend=True)
show(fig)

Interpretation: the points follow the dashed line only loosely (correlation about 0.56). The model gets the direction right: ML Engineer and Data Scientist postings are predicted higher than Data Analyst postings, and they are higher in reality. But within each role the actual pay varies widely, from about $50K to $363K across the whole sample, and the four features cannot explain that spread. Employer, scope of the job, and skills that are not in the model are likely drivers. This is why we describe the model as a guide to relative differences and not as a calculator of someone’s likely salary.

Part 2: Role Opportunity Scorecard (Career Evaluation Logic)

What was scored. Each of the four pathway roles receives a 0 to 1 score on three criteria, combined into one opportunity score.

Why it matters for a job seeker. Pay is only one part of a good opportunity. A role that pays well but asks for senior titles and graduate degrees may not be reachable, and a role that does not match your skills needs more preparation. The scorecard puts the three trade-offs in one table.

Variables and definitions.

Criterion Weight How it is measured Why this weight
Pay 40% Median disclosed salary for the role, divided by the highest role median Pay is the main reason to move along the pathway and is the best-measured of the three
Accessibility 30% Average of two shares: postings whose title is not senior, and postings (with text) that do not require or prefer a Master’s or PhD Shows how realistic entry is for a student or early-career job seeker
Fit with our team 30% Our team’s average self-rating on the six skill areas, weighted by how often the role’s postings name each area, divided by 5 Connects the market to the skills we actually have; see Skill Gap Analysis

The weights are our judgment, not estimated from data, so we test how much the ranking depends on them.

How it was evaluated. A scorecard has no accuracy metric. Instead we check sensitivity: we re-rank the roles under three other sets of weights and report whether the order changes. Each criterion also shows its sample size, because the roles have between 30 and 123 postings with a salary.

team = pd.DataFrame({
    "Python": [5, 3, 4], "SQL": [5, 4, 3], "Machine Learning": [3, 3, 5],
    "Cloud/AWS": [4, 5, 3], "Data Visualization": [4, 3, 5], "Statistics": [2, 4, 3],
}, index=["Manan Patel", "Richard Park", "Michael Phillips"])
team_avg = team.mean()
skill_areas = list(team_avg.index)
roles = ["Data Analyst", "Data / Analytics Engineer", "Data Scientist", "ML Engineer"]

text_posts = mp[mp["has_text"]]
rows = []
for role in roles:
    g, tg = mp[mp["role"] == role], text_posts[text_posts["role"] == role]
    demand = {s: tg["skill_" + s].astype(bool).mean() for s in skill_areas}
    rows.append({
        "Role": role,
        "Postings": len(g), "With salary": int(g["salary_mid"].notna().sum()), "With text": len(tg),
        "Median salary": g["salary_mid"].median(),
        "Not a senior title": (g["level"] == "Other title").mean(),
        "No graduate degree named as required or preferred":
            1 - (tg["degree_Master"].isin(["required", "preferred"]) | tg["degree_PhD"].isin(["required", "preferred"])).mean(),
        "Fit": sum(demand[s] * team_avg[s] for s in skill_areas) / sum(demand.values()) / 5,
    })
sc = pd.DataFrame(rows).set_index("Role")
sc["Pay score"] = sc["Median salary"] / sc["Median salary"].max()
sc["Accessibility score"] = (sc["Not a senior title"] + sc["No graduate degree named as required or preferred"]) / 2
sc["Fit score"] = sc["Fit"]

WEIGHTS = {"Pay score": 0.4, "Accessibility score": 0.3, "Fit score": 0.3}
for k, w in WEIGHTS.items():
    sc["w_" + k] = w * sc[k]
sc["Opportunity score"] = sum(sc["w_" + k] for k in WEIGHTS)
sc = sc.sort_values("Opportunity score", ascending=False)

table = sc[["Postings", "With salary", "With text", "Median salary", "Pay score", "Accessibility score",
            "Fit score", "Opportunity score"]].copy()
table["Median salary"] = table["Median salary"].map("${:,.0f}".format)
table.round(3)
Postings With salary With text Median salary Pay score Accessibility score Fit score Opportunity score
Role
ML Engineer 103 58 102 $220,260 1.000 0.645 0.760 0.822
Data Scientist 211 92 137 $204,875 0.930 0.691 0.746 0.803
Data / Analytics Engineer 335 123 246 $175,000 0.795 0.671 0.786 0.755
Data Analyst 116 30 85 $114,750 0.521 0.788 0.757 0.672
order = sc.index.tolist()[::-1]
fig = go.Figure()
for key, color in [("Pay score", PALETTE[3]), ("Accessibility score", PALETTE[1]), ("Fit score", PALETTE[0])]:
    fig.add_trace(go.Bar(
        y=order, x=[sc.loc[r, "w_" + key] for r in order], orientation="h",
        name=f"{key.replace(' score', '')} ({WEIGHTS[key]:.0%})", marker_color=color,
        customdata=[sc.loc[r, key] for r in order],
        hovertemplate="<b>%{y}</b><br>" + key + ": %{customdata:.2f} (weighted %{x:.2f})<extra></extra>"))
for r in order:
    fig.add_annotation(x=sc.loc[r, "Opportunity score"], y=r, text=f"<b>{sc.loc[r, 'Opportunity score']:.3f}</b>",
                       showarrow=False, xanchor="left", xshift=6)
fig.update_layout(barmode="stack")
fig.update_xaxes(title="Opportunity score (weighted sum of three criteria, maximum 1.0)", range=[0, 0.95])
style(fig, "Role Opportunity Scorecard", "Weights: pay 40%, accessibility 30%, fit with our team 30%", height=470, legend=True)
fig.update_layout(legend=dict(orientation="h", x=0.5, xanchor="center", y=-0.28, yanchor="top"), margin=dict(b=120))
show(fig)
#### sensitivity of the ranking to the weights
scenarios = {
    "Base (40 / 30 / 30)": (0.4, 0.3, 0.3),
    "Equal weights": (1 / 3, 1 / 3, 1 / 3),
    "Pay-heavy (60 / 20 / 20)": (0.6, 0.2, 0.2),
    "Access and fit-heavy (20 / 40 / 40)": (0.2, 0.4, 0.4),
}
ranks = pd.DataFrame({
    name: (w[0] * sc["Pay score"] + w[1] * sc["Accessibility score"] + w[2] * sc["Fit score"])
    .rank(ascending=False).astype(int)
    for name, w in scenarios.items()
}).loc[sc.index]
ranks.index.name = "Role (rank 1 = best)"
ranks
Base (40 / 30 / 30) Equal weights Pay-heavy (60 / 20 / 20) Access and fit-heavy (20 / 40 / 40)
Role (rank 1 = best)
ML Engineer 1 1 1 1
Data Scientist 2 2 2 2
Data / Analytics Engineer 3 3 3 3
Data Analyst 4 4 4 4

Interpretation: ML Engineer (0.822) and Data Scientist (0.803) are the strongest opportunities, ahead of Data / Analytics Engineer (0.755) and Data Analyst (0.672). The order is the same under all four weightings in the table, so the ranking does not depend on our choice of weights, although the gap between the top two (0.019) is small and should not be over-read. The scores show a clear trade-off. ML Engineer has the highest median salary (about $220K) and the lowest accessibility (0.65: 38% of its postings carry a non-senior title, and 9% name a Master’s or PhD as required or preferred). Data Analyst has the lowest median pay (about $115K) but the highest accessibility (0.79: 65% non-senior titles, and 7% name a graduate degree), which fits a student’s first step. Data / Analytics Engineer sits in the middle on both (median about $175K, accessibility 0.67). Data Scientist is close to ML Engineer on pay (about $205K) with slightly better accessibility (0.69, with 47% non-senior titles). Fit with our team is similar across roles (0.75 to 0.79) because our self-ratings are fairly even across the six skills, so fit does not separate the roles; the scorecard’s ranking is driven mostly by pay and accessibility.

What this means for a job seeker. Treat the roles as steps on one path: a Data Analyst role is the most reachable entry point, and the Data Scientist and ML Engineer roles pay nearly twice as much (a median of about $205K to $220K against $115K for analysts) once you have the experience for them. Both are less accessible than the analyst role, so they are the goal of the path, not its starting point. This supports the recommendations in the Skill Gap Analysis page, which identify Machine Learning and Python as the bridge between the first two steps.

Part 3: Salary Estimator

Use the estimator to enter the details of a posting, such as the role, location, experience, degree, and skills, and see what the model predicts. Unlike Parts 1 and 2, it also analyzes your inputs: it shows what is pushing the estimate up or down, which of those drivers are statistically reliable, and which single changes are associated with a higher salary. The analysis runs in your browser from the fitted model below. No outside AI service is used, so nothing you enter leaves the page.

Why a broader training set. Only 303 postings in NAICS 5182 disclose a salary, which is too few to learn from this many inputs (adding skills, degrees, and experience to the Part 1 model barely helped there). So the estimator is trained on all Data Analyst, Data Scientist, ML Engineer, and Data / Analytics Engineer postings from 2026 that disclose a salary, in any industry, from the same Jobs_2026 files. Industry is one of the inputs (“In NAICS 5182” or “Other industry”), so you can see whether our industry pays differently. Posting text is needed to read skills, degrees, and experience, which leaves 2,252 training postings (155 of them in NAICS 5182). The main analysis on every other page stays limited to NAICS 5182.

sm = pd.read_csv("data/processed/met_salary_model_panel.csv")
sm = sm[sm["has_text"]].copy()
sm["salary_mid"] = (sm["salary_min_annual"] + sm["salary_max_annual"]) / 2
# drop a handful of implausible salaries (monthly or typo values)
sm = sm[sm["salary_mid"].between(30000, 500000) & (sm["salary_min_annual"] >= 20000)].copy()

sm["level"] = np.where(sm["title"].str.lower().str.contains(SENIOR), "Senior title", "Other title")
state = sm["state"].fillna("")
top_states = [s for s in state.value_counts().index[:10] if s]
sm["state_group"] = np.where(state.isin(top_states), state, np.where(state == "", "State not given", "Other states"))
sm["remote"] = sm["remote_status"].fillna("Unknown").replace({"Unknown": "Not stated"})
emp = sm["employment_type"].fillna("").str.lower().str.replace(r"[_\-]", " ", regex=True)
sm["emp"] = np.select([emp.str.contains("full"), emp.str.contains("contract"), emp.str.contains("part"), emp == ""],
                      ["Full time", "Contract", "Part time", "Not stated"], "Other")
sm["industry"] = np.where(sm["naics_5182"], "In NAICS 5182", "Other industry")


def highest_degree(row):
    for col, label in [("degree_PhD", "PhD"), ("degree_Master", "Master's"), ("degree_Bachelor", "Bachelor's")]:
        if row[col] in ("required", "preferred", "mentioned"):
            return label
    return "No degree named"


sm["degree"] = sm.apply(highest_degree, axis=1)
sm["exp_known"] = sm["experience_min_years"].notna().astype(int)
sm["exp_years"] = sm["experience_min_years"].fillna(0)
EST_SKILLS = [c[len("skill_"):] for c in sm.columns if c.startswith("skill_")]
for sk_name in EST_SKILLS:
    sm["s_" + sk_name] = sm["skill_" + sk_name].map({True: 1, False: 0, "True": 1, "False": 0}).fillna(0).astype(int)

EST_CATS = ["role", "level", "state_group", "remote", "emp", "industry", "degree"]
EST_NUMS = ["exp_years", "exp_known"] + ["s_" + s_ for s_ in EST_SKILLS]
Xe, ye = sm[EST_CATS + EST_NUMS], sm["salary_mid"].values


def est_pipe(model):
    pre = ColumnTransformer([("cat", OneHotEncoder(handle_unknown="ignore"), EST_CATS), ("num", "passthrough", EST_NUMS)])
    return Pipeline([("prep", pre), ("model", model)])


rows = []
for name, model in {"Baseline (median)": DummyRegressor(strategy="median"),
                    "Ridge regression (all inputs)": Ridge(alpha=10.0)}.items():
    r = cross_validate(est_pipe(model), Xe, ye, cv=RepeatedKFold(n_splits=5, n_repeats=3, random_state=0),
                       scoring=("neg_root_mean_squared_error", "neg_mean_absolute_error", "r2"))
    rows.append({"Model": name, "RMSE (USD)": -r["test_neg_root_mean_squared_error"].mean(),
                 "MAE (USD)": -r["test_neg_mean_absolute_error"].mean(), "R2": r["test_r2"].mean()})
est_results = pd.DataFrame(rows).set_index("Model")
print(f"{len(sm)} training postings, {int(sm['naics_5182'].sum())} of them in NAICS 5182")
est_results.round({"RMSE (USD)": 0, "MAE (USD)": 0, "R2": 2})
2252 training postings, 155 of them in NAICS 5182
RMSE (USD) MAE (USD) R2
Model
Baseline (median) 65656.0 50630.0 -0.01
Ridge regression (all inputs) 43568.0 31996.0 0.56

How well it works. Tested on postings the model had not seen, the estimator’s typical error (MAE) is about $32K and its R² is 0.56, against a typical error of $51K and an R² near zero for guessing the median salary. That is a clear improvement over the Part 1 model (R² 0.31), because it learns from about eight times as many salaries. It is still an estimate: individual postings vary by employer and job scope in ways the inputs cannot capture, so the range shown is wide on purpose.

Salary Estimator and Analyzer

Set the details of a posting. The estimate and the analysis update instantly.
Skills named in the posting

How to read the analysis. The summary line tells you where the estimate falls among all postings the model learned from. “What is driving this estimate” lists the inputs that move it the most compared with an average posting, and keeps only those whose effect stayed clearly above or below zero across 150 resampled refits of the model, so noisy inputs are not presented as findings. “Changes associated with a higher salary” tries one change at a time (another role, a senior title, a different location, a degree, an added skill, or more experience) and lists only the ones whose gain is reliable. These are patterns in postings, so a skill or degree that is associated with higher pay is not guaranteed to raise any one person’s pay. Some effects look counterintuitive. For example, naming Cloud/AWS is associated with slightly lower pay once the role and location are fixed, because skills overlap heavily with roles (an ML Engineer posting almost always names several of them). Read each skill’s effect as “what is left after the role is known”, not as that skill’s value on its own.

Limitations

  • Moderate samples. The regression uses 303 postings with a salary, but only 30 are Data Analyst postings (and 58 are ML Engineer). Differences between close roles, such as ML Engineer and Data Scientist, are not reliable, which is why we report ranges and test the weights.
  • The model explains only part of the variation in pay (R² about 0.31). It shows which factors are associated with pay, not what causes it, and it leaves out experience, employer, and skills.
  • Work arrangement and state estimates are weak. Only 30 postings with a salary are remote, 25 hybrid, and 10 onsite, and several states have few postings, so these effects have wide ranges.
  • The scorecard weights are judgments. The sensitivity check shows which conclusions hold under other weights and which do not.
  • Title-based labels. Role and seniority come from job titles, and a title can overstate or understate the real level of the job. Posting text is available for 570 of the 765 postings, so accessibility rests on fewer postings than pay.
  • Self-ratings. The fit score depends on three people’s self-ratings, which are subjective.
  • The estimator uses a broader dataset. Part 3 is trained on pathway postings from any industry (2,252 postings, 155 of them in NAICS 5182) so that it can use more inputs. Its estimates describe pay in job postings in general, and the industry input shows only how NAICS 5182 compares in that data. The skill, degree, and experience inputs are read from posting text by keyword rules, so they undercount what a posting really asks.