Skip to content
An applied SI research initiativeWrite for the labSupport the lab

Research · 28 Sep 2026

By Siddhant Kasture, Divyaansh Vats, Muhammad Enzi Muzakki, Bishnu Dev Changkakoti

Research noteHQ-S26 note 1

Moderated by HyperQuark Labs · not journal peer review

Status. An HQ-S26 research note: work done by fellows of the HyperQuark Research Fellowship between April and July 2026, written up by the lab from the track's own files. The main result is on synthetic candidates; the test on real candidates was planned and not completed. It has not been reviewed outside HyperQuark, and its code and data have not been released.

The question

Most job-matching counts skills: a candidate who lists more of a role's skills fits better. That misses depth. Two engineers who both list Python, SQL and cloud deployment score the same, whatever each can actually do with them. Standard occupation codes do not help either: one four-digit code covers frontend, backend and machine-learning engineers alike.

The track asked whether a skill graph built from real profile data, with each skill weighted by how much a role needs it, can rank candidates who hold the same skills at different levels, where skill-counting and a standard taxonomy cannot.

HQ-S26 was HyperQuark's first cohort: twelve fellows in four tracks, twelve weeks, weekly reports due every Sunday. No track reached a final submission, so every result here is interim and is reported as the track left it.

Method

Reference graphs. For each role, a skill's weight comes from how often it appears in that role's profiles, grouped into four tiers: Universal (in 90 to 100% of profiles), Strong (60 to 90%), Partial (40 to 60%) and Specialty (20 to 40%). Skills seen in under a fifth of profiles are left out.

Two graphs were built:

  • Machine-learning engineer: 50 public profiles, 24 skills (5 Universal, 1 Strong, 6 Partial, 12 Specialty), total weight 10.99.

  • Software developer: 150 public code-hosting profiles with at least 30 public repositories each, skills text-mined against a fixed 38-skill vocabulary; 24 skills (3, 6, 4 and 11 by tier), total weight 11.42.

The score. With w how much the role needs a skill and u the candidate's proficiency in it (both from 0 to 1), a candidate's fit is the weighted share of the role's needs they meet, counting no skill beyond what the role needs:

fit = Σ min(u_s, w_s) / Σ w_s

The same graph gives a gap list: skills missing or below what the role needs, ordered by w − u.

Baselines. (1) Binary: the share of the graph's skills the candidate has. (2) ESCO: the European skills taxonomy's lists for software developer and, as the closest proxy, artificial-intelligence engineer, essential skills weighted 1 and optional ones 0.5.

Test candidates. Twenty synthetic candidates per role, sixteen after a pilot, in three archetypes: A, every skill present at varied proficiency; B, varied skill sets at fixed proficiency; C, both varied. Because they are synthetic, each candidate's true fit is known. Methods were compared by Spearman's ρ against it, with pass conditions set before the main run and a bootstrap interval (1,000 resamples) on the pooled difference.

A test on real profiles. Forty real profiles were held out (30 software, 10 machine learning). Up to K of each profile's skills were swapped for skills from the other role, five times each, and the methods were scored on how often they still named the right role.

Results

Synthetic candidates (Spearman ρ against the true fit):

RoleArchetype AArchetype BArchetype CPooled (n = 16)
ML engineer: graph score+0.900+0.900+0.886+0.946
ML engineer: binarynot defined+0.800+0.261+0.577
ML engineer: ESCOnot defined+0.975+0.239+0.634
Software developer: graph score+0.900+1.000+0.943+0.976
Software developer: binarynot defined+0.600+0.551+0.747
Software developer: ESCOnot defined+0.564+0.530+0.746

On Archetype A, where every candidate has every skill, binary and ESCO give everyone the same score and cannot rank at all. The pooled gain of the graph score over binary was +0.370 for ML (95% interval +0.054 to +0.829) and +0.229 for software (+0.039 to +0.578). Every pre-registered condition passed on both graphs. Archetype groups held 5 or 6 candidates each.

Real profiles with skills swapped (how often the right role was still named):

Skills swapped (K)01235710
Tier-weighted100%98.5%98.0%92.0%75.5%49.5%9.0%
Unweighted overlap100%99.5%96.0%84.5%52.0%30.5%3.5%

At K = 5 the difference was significant (McNemar, p < 0.001), and every error the weighted method made, the unweighted one made too. When skills were dropped rather than swapped, unweighted overlap was slightly ahead at high K.

Limitations

  • Synthetic candidates carry the main claim, with five or six per archetype. The planned ranking test on real candidates was not reported.

  • The pass conditions were revised after the main run, from pooled conditions to per-archetype ones; the track recorded this openly. Under the original pooled conditions ML passed 5 of 6 and software 3 of 6.

  • Graph stability disagrees across the track's files. Rebuilding the graphs from a fresh sample gave ρ = 0.61 (software) and 0.40 (ML) in the paper draft, and 0.43 and 0.59 in a weekly report. Which is right was not settled.

  • Universal skills dominate. In the ML graph they hold 45.5% of the weight, so a candidate with only the five Universal skills outscores one with all twelve Specialty skills, which is wrong for specialists.

  • Depth above the need is invisible, by design of the minimum; and scores move whenever the graph is rebuilt, so every score must say which graph version made it.

  • ESCO is a weak baseline here: only 8 to 9 of its skills could be matched to the graphs' tool-level vocabulary.

  • No second person has reproduced these results.

What would settle it

The real-candidate ranking test the track designed; the same analysis on a role outside software; and a second person running it from the released code.

Contributions

  • Siddhant Kasture (Track Lead): research direction, tier scheme, software-developer corpus, evaluation harness, synthetic test and swap experiment, ESCO mapping, and the paper drafts this note draws on.

  • Divyaansh Vats: extraction protocol, the machine-learning corpus, the synonym map and the graph-stability runs.

  • Muhammad Enzi Muzakki: tier analysis and threshold sensitivity, the revised tier scheme, the audit of skills left out of the graphs, and simulated reviewer challenges.

  • Bishnu Dev Changkakoti (Programme Director): set the fellowship's research agenda and directed the track.

Data and code

Not released. The profile corpora are described here only in aggregate, and no individual profile is published.

Authors

More research