Data Preparation
Building the Filtered Industry-Career Dataset
Overview
This page documents how we went from the MET Career Compass 2026 job files to data/processed/met_text_panel.csv – the cleaned, filtered dataset behind every analysis page in this product (Exploratory Analysis, Skill Gap Analysis, and Predictive Modeling). Every filtering and cleaning decision below is anchored to our Step 1 scope: the Data Analyst -> Data Scientist / Analytics Engineer career pathway, inside NAICS 5182 (Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services).
Source Data
The source is the Jobs_2026 folder from the course assignments: 20 parquet files holding 815,193 job postings across all industries, all posted between January 2 and September 30, 2026. It is too large for the repository, so we download it beside the repo and filter it with build_met_text_panel.py, which saves only the small result. This is the folder link our instructor confirmed for the project:
gdown --folder "https://drive.google.com/drive/folders/1Tq5Uixwz5J-aG_NfUQdI6X9CIrNz9ZFS?usp=sharing" -O ../Jobs_2026
Only posting fields are copied into the repository. If the instructor republishes the folder, rerunning the script rebuilds the file and the pages update with the new counts.
Industry Filter: NAICS 5182
We keep postings whose three-digit code (NAICS_2022_3) is 518. In these files, all 6,964 postings with that code also carry the four-digit code 5182, so the filter is exactly our selected industry. This removes about 99% of the folder.
Date Window
Every posting in the folder is dated 2026, so no further date filter is needed: the data describes the current market our job seeker faces.
Role Filter: Title Keyword Matching
A posting is on our career pathway if its title contains one of 12 keywords: data scientist, data engineer, machine learning engineer, ml engineer, data analyst, business analyst, bi analyst, business intelligence analyst, analytics engineer, applied scientist, data architect, or database architect. This keeps 1,232 of the 6,964 postings in our industry. We group the result into four role families for the analysis: Data Analyst, Data Scientist, ML Engineer, and Data / Analytics Engineer.
Deduplication
The same job is often listed more than once under different posting IDs (for example, by an employer and a job board). We keep one row per distinct combination of title, company, state, and posting date, which removes 467 repeat listings and leaves 765 postings.
Cleaning the Core Variables
Salary
Salaries are annual. A salary of 0 means not disclosed. A few postings keep a raw hourly rate (under $1,000), which we convert to annual using 2,080 hours. Internship postings are left blank because stipends are not comparable to annual salaries. The salary we analyze is the midpoint of the posted range, and 303 postings (40%) disclose one.
Experience
The files have a minimum-years field, but it is above zero for only 77 postings, so we also read the posting text for phrases such as “5+ years of experience” or “3-5 years’ experience” and take the smallest number tied to experience. Together this gives 206 postings (27%) with a stated minimum.
Skills and Degrees
Two fields in the files are not usable as they are. The skills field holds only about three skills per posting, listed alphabetically, so it cannot rank skills. The education field is filled for every posting, but it is a modeled level and not something the posting states. We therefore searched the posting text (available for 570 of the 765 postings): each skill is matched by a name pattern, and each degree (Bachelor’s, Master’s, PhD) is labeled required, preferred, or named without requirement wording from the sentences that name it. The patterns and rules are in the build script.
Location
The state field mixes full names, abbreviations, and city fragments (such as “Francisco” for San Francisco). We resolve a state from the state code first, then the state-name field, then the location text. This resolves 653 of 765 postings (85%). The rest are labeled Unknown and left out of state charts rather than guessed.
Work Arrangement
Work arrangement is Remote (78 postings), Hybrid (75), Onsite (26), or Unknown (586). About 77% of postings do not state a policy, which is itself a finding.
Company Names
Some employer names are really a web domain (for example, careers.pnc.com) or a job board (“Jobs via Dice”). Known domains are mapped to the real employer, and job-board listings are labeled “Unknown employer”. In total, 27 postings (4%) have no attributable employer name; they are excluded from employer rankings but kept in the data since their role, salary, and location are still valid.
Extended Dataset for the Salary Estimator
The NAICS 5182 slice has only 303 postings with a disclosed salary, which is too few to train a model with many inputs. For the Salary Estimator on the Predictive Modeling page we therefore build a second file from the same Jobs_2026 folder with python build_met_text_panel.py ../Jobs_2026 --all-industries: data/processed/met_salary_model_panel.csv. It keeps the same role keywords, the same repeat-listing removal, and the same salary and text rules, but it keeps pathway postings from any industry and only those that disclose a salary, and it adds a flag for whether the posting is in NAICS 5182.
| Stage | Rows |
|---|---|
| Pathway job-title keywords, any industry | 14,372 |
| One row per distinct posting | 10,713 |
| With a disclosed salary | 2,607 |
| With posting text (used by the estimator) | 2,252 after removing implausible salaries |
Of the 2,252 training postings, 155 are in NAICS 5182. Salaries under $30,000 or above $500,000 are dropped as implausible, and the estimator reads skills, degrees, and experience from the posting text exactly as above. The analysis on every other page uses only the NAICS 5182 file.
Benchmark Tables (All Industries)
The Benchmark Analysis page compares our pathway with the whole job market, so it uses every industry in the same Jobs_2026 folder instead of only NAICS 5182. Because that is 815,193 postings, the script build_benchmark_tables.py reads the folder and saves only two small summary tables, with no posting text or individual rows: benchmark_quantiles.csv and benchmark_fields.csv.
| Stage | Postings |
|---|---|
All postings in the Jobs_2026 files |
815,193 |
| Valid: not a duplicate, possible ghost listing, or internship (using the folder’s own flags) | 733,064 |
| With a usable annual salary (the folder’s normalized pay, $25,000 to $600,000) | 69,246 |
| AI-related (title or text names a core AI term) | about 5% of valid postings |
Occupation fields come from job-title keyword rules (ten computer and business fields, plus “All other occupations”), and a posting is AI-related when its title names AI or machine learning or its text names a core AI term. The rules are listed in the script.
Summary
| Stage | Rows | Notes |
|---|---|---|
All postings in the Jobs_2026 files |
815,193 | All industries and job types, all 2026 |
| NAICS 518 (= 5182) | 6,964 | Our industry |
| Pathway job-title keywords | 1,232 | Excludes unrelated roles |
| One row per distinct posting | 765 | 467 repeat listings removed |
| With posting text | 570 | Used for skills and degree wording |
| With a disclosed salary | 303 | Used for pay analysis and modeling |
| With a stated experience minimum | 206 |
The column-by-column description is in the data dictionary.
See the Exploratory Analysis page for the resulting market baseline.