Psychometrician / IRT–CAT Specialist

Posted 2 weeks ago

Worldwide

Summary

We have built a computer-adaptive reading benchmark and we need someone to tell us where the measurement is wrong. It is a placement and growth instrument, K–12, taken once at intake and then two or three times a year. Item response theory, per-language model selection, EAP ability estimation, Fisher-information item selection, a precision-based stopping rule, exposure control, and a growth model with significance testing. It is calibrated on a large multi-year response corpus. The specification is written for an engineer and runs to about twenty pages. It has been through two independent research passes and an internal review. It has not been read by a psychometrician. That is this job. We want someone who reads the measurement engine and says "that assumption is unsupported, that estimator does the opposite of what you think at the tail, and that bank cannot sustain that exposure policy." We are not looking for validation. We are looking for the things we cannot see. What you would actually do Read the specification and mark it up. All of it, then a second pass at the whole-instrument level. We care about two things: whether each decision is defensible on its own, and whether the decisions are coherent together. The areas we most want examined: Model selection per language, and the assumptions carried across languages Ability estimation and what it does at the extremes of the distribution Item selection, exposure, and whether the bank can support the policy The stopping rule and how the precision target should actually be set Calibration pipeline, parameter recovery, fit and drift Growth measurement, and what is honestly detectable at the individual level The validation design — whether it would survive an independent technical review Deliverable. Comments against numbered sections, plus a short written summary of the instrument-level problems. Direct English, no formatting ceremony. Where you disagree, say what you are disagreeing from — a paper, a standard, or specific experience building or maintaining one of these. Timing. First pass inside two weeks. If the fit is wrong we will both know from the first few sections. Who we are looking for You have built, calibrated or maintained an item bank or an adaptive test. Not read about one. Fluent in IRT parameter estimation (MML/EM or Bayesian), item selection algorithms, exposure control, scale linking or vertical scaling, and DIF. You have opinions about 3PL guessing-parameter stability, about when EAP is the wrong estimator, and about what a small bank does to an exposure policy. Comfortable with the tooling — R (mirt, catR, TAM), flexMIRT, IRTPRO, BILOG, Winsteps, Stata, or equivalent. We do not care which; we care that you have used one in anger. Backgrounds that often fit: psychometrician or measurement scientist at a testing organisation or assessment publisher; educational measurement academic; someone who has run a state or certification testing programme; a statistician who has worked on latent variable models in an assessment context. Who this is not for Machine-learning practitioners who see this as a bandit or recommendation problem. It is a measurement problem with defensibility requirements — a district challenging a child's placement is the design constraint. General data scientists without IRT experience. Reading specialists. We have those. This role is measurement. Anyone who needs the brief re-explained before having a view. Terms Remote, asynchronous, your own hours. NDA before the full specification is shared. Ongoing work available — this is the first of several review passes, and the same person would be our standing measurement reviewer through calibration and validation. THE EXTRACT (Include this verbatim in the posting. Copied unedited from our specification. The screening questions refer to it.) Model, per language. One engine, three maturities. English: 3PL — discrimination, difficulty and guessing all estimated. Hindi: 2PL — difficulty estimated, discrimination borrowed from the English twin item, guessing omitted. Spanish: 1PL — difficulty estimated, discrimination fixed at a = 1, guessing omitted. Language is a data-maturity setting, not a code fork. Ability estimation. EAP (Expected A Posteriori). The prior is set from the child's enrolled grade. Finite and stable at the extremes where MLE diverges, and converges faster from a grade-informed start. Item selection. Maximum Fisher information at the current θ, drawn randomesque from the top 5–10 most informative items. Wrapped in Sympson-Hetter exposure control plus a-stratification. The stopping rule (the one tunable knob). Stop when SE(θ) ≤ 0.30 logits (≈ 0.90 marginal reliability). Floor: at least ~8 items before stopping is allowed. Ceiling: cap at ~25 items regardless, for fatigue. Open decision in our spec: "the precision target is a placeholder… Steven to tune it to a product spec (e.g. place within ½ a grade band, 95% of the time), which maps to a specific SE once the Lexile linking constants are fixed." Section 13 of the same document is titled: "Outputs — the grade scale (no Lexile)." The bank. 288 items, roughly 32 per grade band. Twelve are already flagged for retirement, eight of them because they sit at or below the four-option guess floor. Administered 8–25 items per sitting, two to three sittings per year, per child. Growth. Change between benchmark windows is significance-tested against the combined standard errors of the two estimates, rather than reported as a raw difference.

  • Less than 30 hrs/week
    Hourly
  • < 1 month
    Duration
  • Intermediate
    Experience Level
  • $30.00

    -

    $50.00

    Hourly
  • Remote Job
  • One-time project
    Project Type
Skills and Expertise
Mandatory skills
Social Media Marketing
Nice-to-have skills
Content Writing
Facebook
Activity on this job
  • Proposals:5 to 10
  • Last viewed by client:6 days ago
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Mar 25, 2018
  • United States
    Bellmore9:02 AM
  • $659K total spent
    223 hires, 81 active
  • 19,773 hours
  • Education
    Mid-sized company (10-99 people)

Explore similar jobs on Upwork

Data Analysis
Growth Analytics
Marketing Analytics
Product Analytics
Sales Analytics
Market Analysis
Trade Data Specialist for Flax FiberFixed-price‐ Posted 1 month ago
Microsoft Excel
Data Analysis
Data Mining
Data Modeling

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo