Skill Gap Analysis

Team Skills vs. What Data Analyst and Data Scientist Postings Ask For

Overview

This page compares our team’s current skill levels against how often each skill appears in real job postings for our pathway, Data Analyst to Data Scientist, inside NAICS 5182. It shows where the largest gaps are, how the skills of the two roles differ, and then proposes an improvement plan and a draft set of recommendations for job seekers.

Team Skill Levels

Each team member self-rated their current proficiency on a 1-5 scale:

  • 1 = Beginner
  • 2 = Basic knowledge
  • 3 = Intermediate
  • 4 = Advanced
  • 5 = Expert
import pandas as pd
import plotly.express as px
import plotly.graph_objects as go

PALETTE = ["#457b9d", "#2a9d8f", "#e9c46a", "#e76f51", "#6d597a", "#f4a261"]
GRAY = "#b8bec6"


def style(fig, title, subtitle, height=450):
    fig.update_layout(
        title=dict(text=f"<b>{title}</b><br><span style='font-size:12px;color:gray'>{subtitle}</span>",
                   x=0, xanchor="left"),
        template="simple_white", height=height,
        font=dict(family="Segoe UI, Helvetica, Arial, sans-serif", size=13),
        margin=dict(l=10, r=20, t=95, b=50), hoverlabel=dict(bgcolor="white", font_size=13),
        legend=dict(orientation="h", x=1, xanchor="right", y=1.02, yanchor="bottom"))
    return fig


def show(fig):
    fig.show(config={"displaylogo": False, "responsive": True, "modeBarButtonsToRemove": ["lasso2d", "select2d"]})


skills_data = {
    "Name": ["Manan Patel", "Richard Park", "Michael Phillips"],
    "Python": [5, 3, 4],
    "SQL": [5, 4, 3],
    "Machine Learning": [3, 3, 5],
    "Cloud/AWS": [4, 5, 3],
    "Data Visualization": [4, 3, 5],
    "Statistics": [2, 4, 3],
}

df_skills = pd.DataFrame(skills_data)
df_skills.set_index("Name", inplace=True)
df_skills
Python SQL Machine Learning Cloud/AWS Data Visualization Statistics
Name
Manan Patel 5 5 3 4 4 2
Richard Park 3 4 3 5 3 4
Michael Phillips 4 3 5 3 5 3
fig = go.Figure(go.Heatmap(
    z=df_skills.values, x=list(df_skills.columns), y=list(df_skills.index),
    zmin=1, zmax=5, colorscale="Blues", text=df_skills.values, texttemplate="%{text}", textfont=dict(size=16),
    xgap=3, ygap=3, colorbar=dict(title="Level", tickvals=[1, 2, 3, 4, 5]),
    hovertemplate="<b>%{y}</b><br>%{x}: level %{z}<extra></extra>"))
fig.update_yaxes(autorange="reversed")
style(fig, "Team Skill Levels Heatmap", "Self-rated proficiency, 1 (Beginner) to 5 (Expert)", height=360)
show(fig)

Interpretation: no single team member is strongest across every skill – Manan leads in Python/SQL, Michael leads in Machine Learning/Data Visualization, and Richard leads in Cloud/AWS. Statistics has the lowest team average (3.0), with ratings ranging from 2 to 4. Richard meets the Advanced level; Manan and Michael fall below it.

What the Market Asks For

To ground the comparison in real postings, we used the MET Career Compass 2026 jobs data shared with the course (the Jobs_2026 folder from our assignments, 815,193 postings). We filtered it to our industry and pathway: NAICS 518, then pathway job titles, then one row per distinct posting, giving 765 postings (see Data Preparation). The script is build_met_text_panel.py.

The skills field in these files is truncated (about three skills per posting, listed alphabetically), so it cannot rank skills. Instead we searched the full posting text for each skill. That text is available for 570 of the 765 postings, and those 570 are the base for everything on this page. We grouped postings by role from the job title: Data Analyst (85 postings), Data Scientist or ML Engineer (239, pooled because both are the target of our pathway), and Data / Analytics Engineer (246). A posting counts for a skill when its text matches the skill’s name patterns, which are listed in the build script.

import re

mp = pd.read_csv("data/processed/met_text_panel.csv")
mp = mp[mp["has_text"]].copy()
mp["group"] = mp["role"].map({
    "Data Analyst": "Data Analyst",
    "Data Scientist": "Data Scientist / ML",
    "ML Engineer": "Data Scientist / ML",
    "Data / Analytics Engineer": "Data / Analytics Engineer",
})
groups = ["Data Analyst", "Data Scientist / ML", "Data / Analytics Engineer"]
group_n = mp["group"].value_counts()
skill_cols = [c for c in mp.columns if c.startswith("skill_")]
demand = (mp.groupby("group")[skill_cols].apply(lambda g: g.astype(bool).mean()).T
          .rename(index=lambda c: c[len("skill_"):])[groups])

six = ["Python", "SQL", "Machine Learning", "Cloud/AWS", "Data Visualization", "Statistics"]
heat = demand.loc[six]

fig = go.Figure(go.Heatmap(
    z=heat.values, x=[f"{g}<br>(n={group_n[g]})" for g in groups], y=six, zmin=0, zmax=1, colorscale="Blues",
    text=heat.values, texttemplate="%{text:.0%}", textfont=dict(size=15), xgap=3, ygap=3, showscale=False,
    hovertemplate="<b>%{y}</b> in %{x}<br>Named in %{z:.0%} of postings<extra></extra>"))
fig.update_yaxes(autorange="reversed")
style(fig, "Skill Demand by Role", "Share of postings whose text names each skill area | 570 postings with text", height=420)
show(fig)

Interpretation: the two roles ask for clearly different things. Data Scientist and ML Engineer postings (n=239) name Machine Learning most often (68%), followed by Python (54%), Cloud/AWS (28%), and Statistics (26%). Data Analyst postings (n=85) lean on SQL (31%), Data Visualization (21%), and Statistics (21%), and name Machine Learning in only 11% of cases. Data / Analytics Engineer postings (n=246) emphasize Cloud/AWS (55%), Python (45%), and SQL (43%). These shares are modest overall because posting text is often short and a skill may be written in other words, so read them as relative differences between roles, not as exact requirements. The Data Analyst group (85 postings) is the smallest, so its percentages move more with a few postings.

extra = ["SQL", "Python", "Machine Learning", "Cloud/AWS", "ETL", "Deep Learning", "Data Visualization",
         "Dashboards", "Excel", "Statistics", "Communication", "Spark"]
cmp = demand.loc[extra, ["Data Analyst", "Data Scientist / ML"]]
cmp = cmp.loc[(cmp["Data Scientist / ML"] - cmp["Data Analyst"]).sort_values().index]

fig = go.Figure()
for col, color in [("Data Analyst", PALETTE[0]), ("Data Scientist / ML", PALETTE[3])]:
    fig.add_trace(go.Bar(
        x=cmp[col], y=[f"<b>{s}</b>" if s in six else s for s in cmp.index], orientation="h",
        name=f"{col} (n={group_n[col]})", marker_color=color,
        text=[f"{v:.0%}" for v in cmp[col]], textposition="outside", cliponaxis=False,
        hovertemplate="<b>%{y}</b><br>" + col + ": %{x:.0%} of postings<extra></extra>"))
fig.update_layout(barmode="group")
fig.update_xaxes(title="Share of postings whose text names the skill", tickformat=".0%", range=[0, 0.8],
                 gridcolor="#eceff3")
style(fig, "Skills Requested: Analyst vs. Scientist / ML",
      "Bold = a skill area our team rated | 570 postings with text", height=640)
fig.update_layout(legend=dict(orientation="h", x=0.5, xanchor="center", y=-0.12, yanchor="top"), margin=dict(b=110))
show(fig)

Interpretation: the chart is sorted from the skills most tied to Data Scientist / ML postings (top) to those most tied to Data Analyst postings (bottom). Moving along the pathway swaps a visualization and spreadsheet toolkit (Data Visualization 21%, Dashboards 16%, Excel 18% for analysts, against 8%, 5%, and 1% for scientists) for a modeling and engineering toolkit (Machine Learning 68%, Python 54%, Cloud/AWS 28%, Deep Learning 41%, ETL 20%). SQL is the shared skill: it is the most-named technical skill for analysts (31%) and remains common for scientists (24%).

Comparing Team Skills to Market Demand

The scorecard lists our six skill areas, sorted by how often Data Scientist / ML postings name them, because that is where our pathway leads. We group demand into tiers using the higher of the two role shares (High is 50% or more of postings, Moderate is 25-49%, Low is under 25%) and use Advanced (4) as the level a team member should reach in a high- or moderate-demand skill. The tiers and the target level are team-chosen reading aids; the demand percentages come from the postings.

TARGET = 4


def demand_tier(share):
    if share >= 0.5:
        return "High (50%+)"
    if share >= 0.25:
        return "Moderate (25-49%)"
    return "Low (under 25%)"


order = demand.loc[six].sort_values("Data Scientist / ML", ascending=False).index.tolist()
top_share = demand.loc[order, ["Data Analyst", "Data Scientist / ML"]].max(axis=1)
scorecard = pd.DataFrame({
    "Data Analyst postings": [f"{demand.loc[s, 'Data Analyst']:.0%}" for s in order],
    "Data Scientist / ML postings": [f"{demand.loc[s, 'Data Scientist / ML']:.0%}" for s in order],
    "Demand tier": [demand_tier(top_share[s]) for s in order],
    "Team Average": [round(df_skills[s].mean(), 2) for s in order],
    "Team Minimum": [int(df_skills[s].min()) for s in order],
    "Members below Advanced (4)": [", ".join(df_skills.index[df_skills[s] < TARGET]) or "none" for s in order],
}, index=pd.Index(order, name="Skill"))
scorecard
Data Analyst postings Data Scientist / ML postings Demand tier Team Average Team Minimum Members below Advanced (4)
Skill
Machine Learning 11% 68% High (50%+) 3.67 3 Manan Patel, Richard Park
Python 20% 54% High (50%+) 4.00 3 Richard Park
Cloud/AWS 12% 28% Moderate (25-49%) 4.00 3 Michael Phillips
Statistics 21% 26% Moderate (25-49%) 3.00 2 Manan Patel, Michael Phillips
SQL 31% 24% Moderate (25-49%) 4.00 3 Michael Phillips
Data Visualization 21% 8% Low (under 25%) 4.00 3 Richard Park
rows = scorecard.index.tolist()[::-1]
avg = df_skills.mean()[rows]
low = df_skills.min()[rows]
high = df_skills.max()[rows]

fig = go.Figure(go.Bar(
    x=list(avg), y=rows, orientation="h", marker_color=PALETTE[0],
    text=[f"{v:.1f}" for v in avg], textposition="inside", insidetextanchor="start", textfont=dict(color="white", size=13),
    error_x=dict(type="data", symmetric=False, array=list(high - avg), arrayminus=list(avg - low),
                 color="black", thickness=1.5, width=5),
    customdata=[[lo, hi] for lo, hi in zip(low, high)],
    hovertemplate="<b>%{y}</b><br>Team average: %{x:.1f}<br>Lowest: %{customdata[0]} | Highest: %{customdata[1]}<extra></extra>"))
fig.add_vline(x=TARGET, line_dash="dash", line_color=PALETTE[3], line_width=3, layer="above",
              annotation_text="Target: Advanced (4)", annotation_position="top")
fig.update_xaxes(title="Self-rated proficiency (1-5)", range=[0, 5.2], gridcolor="#eceff3")
style(fig, "Team Average vs. Target by Skill", "Sorted by Data Scientist / ML demand | black lines show the lowest to highest member rating", height=460)
show(fig)

Interpretation: Machine Learning is the most-requested skill for Data Scientist / ML postings (68%) and one of the team’s weaker areas: the team averages 3.7, and Manan and Richard each rate 3, so it is the clearest gap for our pathway. Python (54% of Data Scientist / ML postings) is also in the high-demand tier; the team averages 4.0, but Richard rates it 3. Cloud/AWS (28%) is in the moderate tier, with the team at 4.0 and Michael at 3. SQL (31% of Data Analyst postings) and Statistics matter for the analyst step; the team averages 4.0 in SQL, with Michael at 3. Statistics is the team’s lowest average (3.0, with Manan at 2 and Michael at 3), and it appears in 26% of Data Scientist / ML and 21% of Data Analyst postings, so it is a real gap and not only a self-rating quirk. Data Visualization is named in only 8% of Data Scientist / ML postings (21% of Data Analyst postings), so it matters mostly for the first step of the path.

Improvement Plan

These proposed milestones should be confirmed by the team, and each suggested resource should be checked for current availability. Timeframes begin when the plan is adopted; they describe future work, not completed contributions.

Member Priority (self-rating; share of Data Scientist / ML or Data Analyst postings) Proposed action Suggested resource Timeframe and evidence
Manan Patel Machine Learning (3; 68% of Data Scientist / ML), then Statistics (2; 26%) Pair with Michael on a small supervised-learning project; practice hypothesis testing and regression with Richard’s review. “Machine Learning Specialization” (Coursera, DeepLearning.AI); Khan Academy, “Statistics and Probability” (free) Within two weeks: one reproducible notebook explaining assumptions, results, and limitations.
Richard Park Machine Learning (3; 68%), Python (3; 54%), Data Visualization (3; 21% of Data Analyst) Pair with Michael on Machine Learning and visualization and with Manan on Python. “Python for Everybody” (Coursera, University of Michigan); Tableau Public free training videos Within two weeks: one Python analysis with two labeled charts and written interpretations.
Michael Phillips SQL (3; 31% of Data Analyst), then Cloud/AWS (3; 28% of Data Scientist / ML) and Statistics (3; 26%) Start with SQL aggregation practice, then a small AWS analysis because of the industry scope. Mode, “SQL Tutorial” (free); “AWS Cloud Practitioner Essentials” (AWS Skill Builder) Within two weeks: SQL aggregation queries with written results; within four weeks: a small analysis run on AWS, documenting setup, outputs, and resource cleanup.

Shared priority: Machine Learning is the most-requested skill for the role our pathway leads to, and Michael (rating 5) can coach Manan and Richard. Richard and Michael can also teach each other’s weaker skill (Richard rates SQL at 4 and Michael rates Data Visualization at 5). Statistics remains the team’s lowest average, so within two weeks the team should also complete one worked statistical example and peer-review its interpretation.

Draft Recommendations

These draft recommendations for job seekers come from the Exploratory Analysis page and the market comparison above. They describe the Data Analyst to Data Scientist pathway in NAICS 5182.

  1. Learn the analyst toolkit first, then add modeling. Data Analyst postings most often name SQL (31%) and Data Visualization (21%), with Excel (18%) and Dashboards (16%) close behind. Build portfolio dashboards backed by SQL queries before moving on.
  2. Make Machine Learning and Python the bridge to Data Scientist roles. Data Scientist / ML postings name Machine Learning in 68% of cases and Python in 54%, compared with 11% and 20% for Data Analyst postings.
  3. Add Cloud/AWS and data-pipeline skills. They appear in 28% and 20% of Data Scientist / ML postings and far more often in engineering roles (55% and 68%), which fits an industry built around cloud and data hosting.
  4. Expect a mid-level market and build experience early. Where experience is stated, requirements mostly fall in the 3-8 year range, so projects, internships, and certifications can help offset limited work history.
  5. Look at California and New York for volume, and California for pay. California and New York have the most postings (141 and 113), and California leads on average disclosed pay ($225K across 85 salary postings).
  6. Weigh the pay gain. Average disclosed pay for Data Scientist and ML Engineer postings is about 1.8 times that of Data Analyst postings (about $219K against $125K). The Predictive Modeling page finds role to be the strongest pay factor.
  7. Do not assume a degree is stated. Only a minority of postings name a degree; see the education section of the Exploratory Analysis page for how many state it as required.

Limitations and Next Steps

  • Posting text is available for 570 of the 765 postings, and some of that text is short, so skill shares are understated and should be compared across roles, not read as exact requirements.
  • The Data Analyst group is the smallest (85 postings with text, and only 30 of all 116 Data Analyst postings disclose a salary), so role comparisons are directional.
  • Skills are matched by name patterns (listed in build_met_text_panel.py), so a posting that names a skill in another way is missed, and a pattern can match unrelated uses of a word.
  • The target level (Advanced, 4) is a team choice, not a figure taken from postings.
  • Self-ratings are subjective. Each member should confirm their ratings, particularly any 5 (Expert), with a project or work example.
  • Review the proposed milestones together and reassess individual gaps after the deliverables are completed.