The Skills Graph
1,533 atomic skills, cross-linked to 2,536 careers, refreshed monthly from live job-posting evidence and anchored to O*NET and ESCO. The substrate for skill-gap analysis and the targeted free-course recommendations that follow.
What is in the skills graph
A node is a skill: an atomic competency that can be learned, practiced, and measured independently of other competencies at a meaningful level. There are 1,533 of these nodes today. An edge is a relationship between a skill and a career, weighted by importance or empirical frequency and tagged with its provenance. Skill nodes also carry secondary edges among themselves — prerequisite edges (you typically need linear algebra before deep learning), substitute edges (Postgres and MySQL fill the same role in many career contexts), and parent–child edges (back-end web development is a parent of Node.js, Django, FastAPI).
Where the skill nodes come from
O*NET Content Model.The U.S. Department of Labor's occupational database recognizes ~35 occupational skills per role, plus knowledge areas, abilities, and detailed work activities. O*NET is the anchor for the long-stable skill universe — communication, problem solving, manual dexterity, quantitative reasoning — and the canonical importance ratings live here.
ESCO Skills Pillar. The European framework ESCO maintains a finer-grained skills pillar with ~13,000 skill and competence nodes — particularly strong on software, foreign languages, and specific technical methods. We import targeted subsets of ESCO to fill the gaps where O*NET aggregates skills at a higher level than is useful for course recommendation.
Live posting evidence.Genuinely new skills — prompt engineering, LLM evaluation, retrieval-augmented generation, carbon-accounting frameworks — appear in job postings months or years before they enter O*NET. We ingest de-identified postings, canonicalize skill mentions, and admit a new node when posting frequency crosses an empirically calibrated threshold. The new node carries an "emerging" tag until the next O*NET refresh either confirms it or supersedes it.
How the skill–career edges are weighted
Same evidence regime as the broader knowledge graph: an edge enters via O*NET importance ≥ 3.0, posting co-occurrence ≥ 15%, or expert review. The weight on the edge is the importance score (1–5 scale) when admitted via O*NET, the empirical frequency (0–1) when admitted via posting evidence, or the expert-assigned importance when admitted via review. Provenance is tagged on every edge so weights from different sources are not silently averaged into a misleading single number.
What skill-gap analysis actually does
A skill-gap analysis answers the question: between where I am and the typical skill profile of the career I want, what should I learn next? The pipeline has three steps.
- Current profile. For each skill the user has data on (assessed via in-product tests, self-rated, or imported from credentials), we know their current level on a normalized scale.
- Target profile.For each career the user is targeting, the graph supplies the typical-incumbent profile — the median required level for each skill weighted by that skill's importance for the career.
- Gap ranking.The per-skill difference is multiplied by the skill's importance for the career and ranked. The top gaps are the skills with the largest weighted distance — not the skills the user is worst at in absolute terms, but the skills whose under-mastery would matter most for the target career.
How course recommendation closes the loop
For every top-ranked skill gap, the system surfaces the highest-quality free or low-cost course we can verify. The catalogue is drawn from a curated list of open-courseware sources — MIT OpenCourseWare, Stanford Online, Coursera open offerings, edX archived offerings, freeCodeCamp, official technology documentation, and a smaller set of audited paid offerings where no open alternative exists. Quality is ranked by independent signals: instructor credentials, syllabus completeness, learner-reported outcomes, and recency.
We do not earn affiliate revenue from course providers. Recommendations sort on independent quality signals, not on referral commissions — a hard-line product decision that costs us a revenue channel competitors do take. The Future of Jobs Report (WEF, 2023) and McKinsey's education-to-employment work (Mourshed et al., 2012) consistently document that mis-aligned course recommendations are a primary cause of skill-program waste; we choose not to add to that pile by introducing financial conflicts.
Refresh cadence and provenance
- O*NET-anchored edges: refreshed annually on the O*NET release schedule (autumn).
- ESCO-derived nodes and edges: refreshed quarterly against ESCO version updates.
- Posting-evidence weights: refreshed monthly with a rolling 90-day window of postings.
- Expert-review edges: audited quarterly against the latest O*NET and posting evidence.
- Course catalogue: recrawled monthly; broken or archived links downgrade or remove the recommendation.
Honest limitations
The most important limitation: the graph reflects what employers ask for, which is a systematically biased view of what jobs require. Postings overstate (a back-end role asking for ten technologies the actual job uses three of) and understate (a senior role whose skills section hides under "you'll know what you need"). Weight aggregation across many postings dampens both biases, but the underlying noise is real and we say so. A skill-gap analysis is a prioritized hypothesis about what to learn, not a contract that learning it will land the job.
The second limitation is course quality. Open-courseware quality varies widely, and our quality ranking is imperfect. We invite users to flag specific recommendations that did not deliver, and flagged recommendations downrank quickly.
The third limitation is geography. Posting evidence is collected predominantly from English-speaking labor markets. Skill weights for careers in non-English markets are derived from a smaller posting base and a heavier reliance on ESCO; this is disclosed on the career page's data-snapshot stamp.
Citations
- U.S. Department of Labor / Employment and Training Administration (2024). O*NET Content Model — Skills, Knowledge, Abilities. https://www.onetcenter.org/content.html link
- European Commission (2022). ESCO Skills Pillar — European Skills, Competences, Qualifications and Occupations. ESCO v1.1.1 skills pillar (~13,000 skills/competences). link
- World Economic Forum (2023). Future of Jobs Report 2023. World Economic Forum, Geneva. link
- Mourshed, M., Farrell, D., & Barton, D. (2012). Education to employment: Designing a system that works. McKinsey Center for Government. link
Three sources combined. O*NET supplies the canonical skill nodes — the platform recognizes 35 occupational skills plus knowledge areas, abilities, and work activities, which together cover the long-stable competencies. ESCO supplies finer-grained tooling skills — particularly software, languages, and technical methods — that O*NET represents at a higher level of abstraction. Live job postings supply emerging skills (LLM prompting, prompt engineering, climate-scenario modeling, others) that have not yet been codified by either authority. The 1,533 number is the de-duplicated union after taxonomy alignment.
A skill is treated as atomic when it can be learned, practiced, and measured independently of others at a meaningful level. "Python" is atomic. "Backend engineering" is not — it is a cluster of atomic skills (Python, SQL, system design, Linux, observability). We err toward atomicity because course recommendation, skill-gap analysis, and credential mapping all work better at the atomic level. Clusters are computed from atomic skills, never the reverse.
Through the same evidence gates that govern the broader knowledge graph: a skill links to a career if O*NET rates its importance at ≥ 3.0 for that occupation, or if it co-occurs in ≥ 15% of recent postings for that occupation, or if expert review admits it for an emerging role. The link carries a weight (the importance or the empirical frequency) and a provenance tag. The same skill can connect to many careers with different weights — Python links to 47 careers in our graph with weights ranging from 2.1 to 4.8 on the O*NET importance scale.
When a user completes a skills assessment or self-rates against a career's skill requirements, the gap between current skill level and required skill level is computed per atomic skill. The largest gaps are surfaced as "skills to close." For each gap, we recommend the highest-rated free or low-cost course we can verify, drawn from a vetted catalogue of open-courseware sources (MIT OCW, Stanford Online, Coursera open courses, edX archived offerings, freeCodeCamp, official documentation). We do not earn affiliate revenue from course providers — recommendations are sorted by independent quality signals, not by referral commissions.
Two reasons. First, auditability: a curated graph lets us trace any skill–career link back to a specific evidence source, which an LLM extraction does not. Second, latency and cost: serving 2,536 career pages against a live LLM extraction at request time is expensive and slow; precomputing the graph keeps pages fast and free for the user. We do use LLM-assisted skill normalization during the ingestion pipeline (mapping "writing JS" and "JavaScript development" to the same canonical node), but the resulting edges are reviewed before they enter the published graph.
Through the posting-evidence gate. When a new skill name appears in job postings at material frequency, the ingestion pipeline canonicalizes it (deduplicating against existing nodes), evaluates whether existing edges should redirect to the new node or coexist, and either auto-admits it (if the evidence is unambiguous) or queues it for expert review. The same pipeline handles fading skills: when posting evidence for a skill drops below the threshold for an extended window, edges decay before the skill is archived. The graph is supposed to reflect today's labor market, not 2014's.
Yes, in two ways we disclose. First, the graph reflects what employers ask for, not what is sufficient for the work — a job posting may overstate (a Python role asking for ten technologies the actual job uses three of) or understate (a senior role with a short skills section). We mitigate this with weight aggregation across many postings, but the bias is real. Second, course quality on the open web is highly variable. We curate the source list, but a recommended course can still be outdated or weakly aligned with the target skill. We invite users to flag recommendations that are not useful so the catalogue improves.