Research and experiential learning for public service
Technical documentation

INSIGHT+ data and codebook.

A high-level guide to sources, transformations, fields, linkages, access, and the quality controls behind the platform.

Architecture

Built to connect records that were never designed to work together.

INSIGHT+ uses an extract–transform–load pipeline to harmonize federal personnel and hiring data with occupational, economic, geographic, and institutional measures.

Raw sources retain their provenance. Standardized identifiers and derived variables are documented so that users can distinguish reported fields from transformations, classifications, model outputs, and linked context.

Documentation status. This prototype describes the platform at a public-facing level. Versioned schema files, source-specific data dictionaries, refresh dates, and disclosure rules should be attached to the production GSU deployment.
Core principles

Reproducible by design

  • Preserve source provenance and extraction dates.
  • Separate raw, standardized, linked, and analytic layers.
  • Use durable agency, occupation, location, and time identifiers.
  • Document machine-assisted classifications and validation.
  • Apply access controls appropriate to each source.
Source families

What the platform brings together.

Specific variables and periods vary by source and access level. The production codebook should record coverage, refresh cadence, restrictions, and transformations for every table.

Source familyRepresentative contentPrimary usesTypical unit
Federal personnelAgency, occupation, grade, duty station, appointment, accession, separation, demographic and career attributesWorkforce stocks, flows, mobility, retention, compositionEmployee-event or employee-period
Federal job postingsAnnouncement text, occupation, agency, location, salary, work schedule, remote/telework language, dates and outcomesRecruitment demand, applicant response, skills, remote workAnnouncement or opportunity
Occupational systemsFederal series, SOC crosswalks, tasks, knowledge, skills and abilitiesCross-sector matching, task analysis, AI modelingOccupation or task
Labor-market contextEmployment, wages, sector comparisons, local and occupational conditionsRecruitment competition, wage gaps, stability premiumOccupation-place-period
Institutional measuresAgency structure, independence, ideology, shutdown exposure, political contextPolicy evaluation, heterogeneity, institutional explanationAgency-period
Derived classificationsJob-text labels, task clusters, AI CAS scores, harmonized locations and occupationsSearch, comparison, prediction, scenario analysisRecord, occupation or task
Representative fields

A compact public codebook.

Names below illustrate a consistent semantic layer; they are not a substitute for the version-specific machine-readable schema.

FieldDescriptionTypeNotes
agency_idStable identifier for agency or organizational unitStringCrosswalked across source naming conventions and reorganizations
occupation_idFederal occupational series with crosswalk to broader classificationsStringCan be linked to SOC and task/skill systems
duty_stationStandardized work locationGeographyMay include city, county, state, commuting zone and metro status
event_dateDate associated with an accession, separation, posting or other eventDateSupports monthly, quarterly and annual aggregation
remote_statusStructured classification of remote or telework eligibilityCategoryDerived from structured fields and announcement text; versioned
shutdown_exposureAgency-level treatment or intensity measure for a shutdown periodNumeric / categoryDefinition depends on study design and data source
cas_complementarityEstimated degree to which AI capabilities complement an occupation1–5 scoreModel version, prompt, retrieval corpus and justification retained
cas_augmentationEstimated potential for AI to improve or extend human task performance1–5 scoreInterpreted jointly with complementarity and substitutivity
cas_substitutivityEstimated potential for AI to perform tasks with reduced human input1–5 scoreScenario indicator, not a deterministic forecast of job loss
Quality and governance

Analytical value depends on disciplined stewardship.

The platform should make uncertainty, model dependence, missingness, revisions, and access constraints visible rather than hiding them behind a dashboard.

Validation

Automated range, duplication, consistency, and referential-integrity checks are paired with source-level spot checks and study-specific validation.

Version control

Data snapshots, transformation code, schema versions, and model outputs should be tagged so published analyses can be reproduced.

Disclosure and access

Public aggregates, restricted microdata, partner data, and licensed sources require distinct access and disclosure rules.

Responsible modeling

Machine-learning and LLM-assisted fields retain provenance, model information, confidence or rationale, and human-review procedures.

Data access

Request a research conversation.

Describe the question, needed unit of analysis, time period, intended use, collaborators, and any disclosure constraints. Availability depends on source permissions and project fit.