Validation
Automated range, duplication, consistency, and referential-integrity checks are paired with source-level spot checks and study-specific validation.
A high-level guide to sources, transformations, fields, linkages, access, and the quality controls behind the platform.
INSIGHT+ uses an extract–transform–load pipeline to harmonize federal personnel and hiring data with occupational, economic, geographic, and institutional measures.
Raw sources retain their provenance. Standardized identifiers and derived variables are documented so that users can distinguish reported fields from transformations, classifications, model outputs, and linked context.
Specific variables and periods vary by source and access level. The production codebook should record coverage, refresh cadence, restrictions, and transformations for every table.
| Source family | Representative content | Primary uses | Typical unit |
|---|---|---|---|
| Federal personnel | Agency, occupation, grade, duty station, appointment, accession, separation, demographic and career attributes | Workforce stocks, flows, mobility, retention, composition | Employee-event or employee-period |
| Federal job postings | Announcement text, occupation, agency, location, salary, work schedule, remote/telework language, dates and outcomes | Recruitment demand, applicant response, skills, remote work | Announcement or opportunity |
| Occupational systems | Federal series, SOC crosswalks, tasks, knowledge, skills and abilities | Cross-sector matching, task analysis, AI modeling | Occupation or task |
| Labor-market context | Employment, wages, sector comparisons, local and occupational conditions | Recruitment competition, wage gaps, stability premium | Occupation-place-period |
| Institutional measures | Agency structure, independence, ideology, shutdown exposure, political context | Policy evaluation, heterogeneity, institutional explanation | Agency-period |
| Derived classifications | Job-text labels, task clusters, AI CAS scores, harmonized locations and occupations | Search, comparison, prediction, scenario analysis | Record, occupation or task |
Names below illustrate a consistent semantic layer; they are not a substitute for the version-specific machine-readable schema.
| Field | Description | Type | Notes |
|---|---|---|---|
agency_id | Stable identifier for agency or organizational unit | String | Crosswalked across source naming conventions and reorganizations |
occupation_id | Federal occupational series with crosswalk to broader classifications | String | Can be linked to SOC and task/skill systems |
duty_station | Standardized work location | Geography | May include city, county, state, commuting zone and metro status |
event_date | Date associated with an accession, separation, posting or other event | Date | Supports monthly, quarterly and annual aggregation |
remote_status | Structured classification of remote or telework eligibility | Category | Derived from structured fields and announcement text; versioned |
shutdown_exposure | Agency-level treatment or intensity measure for a shutdown period | Numeric / category | Definition depends on study design and data source |
cas_complementarity | Estimated degree to which AI capabilities complement an occupation | 1–5 score | Model version, prompt, retrieval corpus and justification retained |
cas_augmentation | Estimated potential for AI to improve or extend human task performance | 1–5 score | Interpreted jointly with complementarity and substitutivity |
cas_substitutivity | Estimated potential for AI to perform tasks with reduced human input | 1–5 score | Scenario indicator, not a deterministic forecast of job loss |
The platform should make uncertainty, model dependence, missingness, revisions, and access constraints visible rather than hiding them behind a dashboard.
Automated range, duplication, consistency, and referential-integrity checks are paired with source-level spot checks and study-specific validation.
Data snapshots, transformation code, schema versions, and model outputs should be tagged so published analyses can be reproduced.
Public aggregates, restricted microdata, partner data, and licensed sources require distinct access and disclosure rules.
Machine-learning and LLM-assisted fields retain provenance, model information, confidence or rationale, and human-review procedures.
Describe the question, needed unit of analysis, time period, intended use, collaborators, and any disclosure constraints. Availability depends on source permissions and project fit.