Event-Study Econometrician — Historical News & Equity Price Reaction Analysis

Posted 2 days ago

Only freelancers located in the U.S. may apply.U.S. located freelancers only

Summary

Engagement: Contract, phased. Phase 1 is a paid, standalone feasibility engagement with a go/no-go gate before Phases 2 and 3 are commissioned. Duration: Phase 1 approx. 1–2 weeks. Full scope approx. 1–3 months. Your compensation and evaluation are tied to the rigor of the method, not to what the data shows. Wherever the evidence is insufficient to support a claim, documenting that is a required and fully compensated deliverable. Our product architecture already assumes selective output: we display a result only where the evidence supports one and return "insufficient evidence" everywhere else. Both halves of that map are load-bearing, and we have no commercial preference for how the cohorts divide between them. We are not looking for a trading signal, an alpha model, or a predictive engine, and work that drifts in that direction is outside the scope of this engagement. What we want is a rigorous, classical event study: given this category of news event and this type of stock, what is the historical distribution of abnormal returns, with honest confidence intervals, and does it hold out of sample? We would rather fund a well-specified analysis whose answer surprises us than a poorly specified one that confirms anything. About Us We are a news curation and context firm serving US equity markets. We deliver curated, context-enriched headlines to over 200 institutions and more than 3,000 institutional traders on both the buy and sell side. We have been doing this for decades and are among the fastest on the street. Two features of our data matter for this work: Our archive is curated, not scraped. Roughly 2,000 items per day pass an editorial filter applied by former traders, PMs, and analysts. The same editorial logic produced our historical archive and our live feed, so there is no train/live mismatch in the sample definition. Timestamps are reliable. We are frequently first to print. Delivery timestamps closely track information arrival. Every headline in the archive carries an AI-generated sentiment label, a directional strength score, an accompanying model confidence score, and classification into up to 25 taxonomies. These labels are inputs to your analysis. Assessing their stability over time is in scope; generating new ones is not. We are building a retail-facing product on top of this archive. Your work establishes its empirical foundation. The Core Question For a given cohort (defined by some combination of stock characteristics, news taxonomy, and sentiment tier) what is the historical distribution of abnormal returns over six horizons following a headline, and how much of that distribution is statistically distinguishable from the cohort's unconditional baseline? The six horizons: Horizon Window 1 ~60 minutes post-headline 2 ~120 minutes post-headline 3 End of day (4:00 PM ET close) 4 Overnight / session gap 5 3 trading days 6 5 trading days Horizon definitions must be measured in trading time, not calendar time, with defined handling for headlines crossing outside regular session hours. Scope of Work Phase 1 : Methodology and Feasibility (paid, standalone, go/no-go) Deliverable: a written methodology document sufficient for a second statistician to reproduce your approach, plus a feasibility assessment against a data sample we provide. Must address: Abnormal return specification. Definition of the market/sector model used to isolate idiosyncratic return. Raw returns and own-volatility normalization alone are not sufficient. Volatility baselines. Time-of-day-conditional volatility estimation for intraday horizons. A trailing daily sigma scaled to one hour will not do. Cohort structure. How to group and condition. Our 25 taxonomies were built for editorial routing, not statistical pooling; some likely merge, some likely split. Our team will work with you on this as we know the market behavior and expect to inform the groupings. Thin-cell handling. Ticker-level cohorts will often have very few observations while pooled cohorts have many thousands. We expect a partial-pooling / hierarchical or empirical-Bayes approach, or a reasoned argument for a different treatment. Minimum evidence threshold. Derive our display threshold as a function of a stated minimum detectable effect and target power. We want the reasoning, not just a number. Dependence structure. Treatment of overlapping estimation windows, multiple headlines on the same name in the same window, and cross-sectional correlation on macro days. Baseline construction. Every cohort needs its own unconditional baseline. Absolute hit rates are not meaningful on their own. Validation design. How you will establish that estimates hold out of sample. Phase 2 : Estimation and Validation Full historical estimation across all valid cohort × taxonomy × sentiment × horizon permutations. Walk-forward out-of-sample validation. Freeze the estimates on an early period, run forward on data you have not touched, and report how many qualifying cohorts held up and by how much the effect decayed. This is a required deliverable and a primary measure of the project's success. Multiple-comparison control appropriate to the number of cells tested. Deliverable: a static reference matrix (CSV or Parquet) containing every cohort, its point estimate, confidence bounds, sample size, unconditional baseline, and qualification status across all six horizons. Deliverable: an explicit register of cohorts where no reliable effect was found. Phase 3 : Logic Transfer Documented mathematical logic and a reference implementation (Python) sufficient for our engineering team to reproduce the calculations in production. The normalization scheme converting validated estimates into a bounded 0–100 display score. Our current thinking is that this should derive from the lower confidence bound on the effect over baseline, so that sample size enters through the statistics rather than as a hand-set weight — we welcome a better proposal. Working session with our engineering team and written sign-off. -Data and Tools- All data is provided by us. No data sourcing, licensing, or acquisition is required on your part. Historical curated headline archive with AI-generated sentiment, strength, confidence, and taxonomy labels. Historical pricing data , which we supply: minute and 5-minute OHLC bars for intraday windows, daily aggregates for multi-day windows, and participant timestamp data covering pre-market (4:00 AM – 9:30 AM ET) and post-market (4:00 PM – 8:00 PM ET) sessions. Extraction volume and format can be adapted to your requirements. Tell us what you need and how you want it delivered. Language and libraries are your choice. Python or R both fine. Methodological Constraints All methods must be classical, auditable, and explainable line by line. This is a retail-facing product; every number we display must be defensible to a regulator, a journalist, or a skeptical user. No black-box or machine-learning models in the estimation pipeline. No neural networks, gradient boosting, or ensemble learners. No LLM-assisted analysis or interpretation. To be unambiguous: standard econometric and statistical tooling is expected and welcome. Hierarchical and random-effects models, empirical Bayes, shrinkage and regularization, bootstrapping, and power analysis are all classical methods and all in scope. The constraint is on opacity, not on sophistication. Explicitly Out of Scope Production software development, database design, or streaming pipeline work. Any form of signal optimization, strategy backtesting, or portfolio construction. Generating or revising sentiment/taxonomy labels. Product, UI, or commercial strategy. Who We're Looking For Strong fit: Formal training in econometrics, statistics, or empirical finance. Direct experience with event-study methodology — abnormal returns, estimation windows, cross-sectional aggregation, the Brown & Warner / MacKinlay / Kothari & Warner literature. Published or applied work measuring the price impact of information events. Comfortable delivering a null result and defending it. Willing to tell us when part of our framing is wrong. Our team is made up of former buy-side and sell-side traders, PMs, and analysts and we know the market, we do not claim to be statisticians, and we want to be pushed back on. Poor fit: Quantitative researchers whose instinct is to iterate until something profitable appears. Anyone who would treat "no effect found" as a problem to be engineered around. Generalist data scientists whose primary toolkit is ML. Anyone who would execute the methodology above without questioning any of it. Academic researchers, finance PhD candidates and postdocs, and practicing econometricians consulting on the side are all encouraged to apply. Screening Task (Paid) We will shortlist a small number of applicants and pay each for a 2–4 hour written task at their stated rate before any larger commitment. The task: You have approximately 5,000 news events across approximately 800 US equities. Propose a methodology for estimating the probability of a directional 3-day abnormal return, conditional on event category and stock characteristics, where per-ticker event counts range from 3 to 400. Address: how you define abnormal return; how you handle thin cells; how you set a minimum evidence threshold for reporting a result; and how you validate that estimates hold out of sample. Two pages. No code required. We are evaluating reasoning and judgment, not polish or length. To Apply Please include: A brief description of the most relevant event study or abnormal-return analysis you have conducted, including what the data showed and what you concluded. One example of a project where your conclusion was that the effect was not there. What you did, and how you presented it to whoever commissioned the work. Your view on how to handle cohorts with very few observations, in two or three sentences. Any published work or writing samples. Availability and rate. Applications that restate this posting back to us will not be reviewed. We would rather see one paragraph of real disagreement with our approach than three pages of agreement. An NDA will be required before data access.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $80.00

    -

    $120.00

    Hourly
  • Remote Job
  • Complex project
    Project Type
Skills and Expertise
Mandatory skills
Statistics
Statistical Analysis
Activity on this job
  • Proposals:Less than 5
  • Last viewed by client:yesterday
  • Hires:
    1
  • Interviewing:
    2
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Mar 19, 2021
  • United States
    Wilton7:18 PM
  • $8.1K total spent
    8 hires, 3 active
  • 515 hours
  • Mid-sized company (10-99 people)

Explore similar jobs on Upwork

Data Analysis
Growth Analytics
Marketing Analytics
Product Analytics
Sales Analytics
Market Analysis

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo