World-record model

Custom neural networks · Financial intelligence

What 100% on OfficeQA certifies for One Jump Inc.

Capability report · Financial-document analysis · Statistical reasoning · Specified-model forecasting

OfficeQA — Full (246 questions) and Pro (133 hard questions)
Corpus — 89K+ U.S. Treasury Bulletin documents

Executive summary

OfficeQA tests considerably more than ARIMA.

OfficeQA combines document retrieval, financial calculations, statistical analysis, and several forecasting methods across 246 questions answered from a corpus of 89K+ U.S. Treasury Bulletin documents.

The methods are prescribed; the evidence is not.

Questions specify the model — often its parameters — and check the final number against a reference. Finding the right tables, periods, and revisions is the agent's job, and where most failures happen.

62% of Pro questions go beyond basic arithmetic.

Regression, smoothing, HP filtering, correlation, VaR, expected shortfall, inflation and currency adjustment, and forecast-versus-actual comparisons all appear in the question set.

Correct execution is not the same as real-world prediction.

A perfect score establishes document-based quantitative analysis and specified forecasts. It does not, by itself, establish model selection, calibrated intervals, or company revenue forecasting.

The accurate description of the result

Financial-document analysis, statistical reasoning, and specified-model forecasting — not merely ARIMA forecasting, and not a complete test of autonomous corporate financial planning.

1The benchmark, by the numbers

OfficeQA Full and Pro retrieve from the same collection of 89K+ Treasury Bulletin documents.

246

questions in OfficeQA Full — all answered correctly

133 hard113 easy
133

hard questions make up OfficeQA Pro — same corpus, easy items excluded

62%

of Pro questions require analysis beyond basic arithmetic (technical paper)

89K+

Treasury Bulletin documents to retrieve from — hierarchical tables, fiscal years, revised figures

2Forecasting methods actually tested

Concrete questions from the benchmark — not methods an agent could theoretically use. Each prescribes the model, and often its parameters; the answer is checked against a reference value with numerical tolerance.

MethodWhat the question asks the agent to doQuestionAppears in
Linear regressionrevenueFit federal individual income-tax receipts from 1929–1942, then extrapolate receipts to a later fiscal year.UID0014Full
Linear regressionspendingUse Agriculture Department spending from 1990–1998 to project 1999 spending, returning the slope, intercept, and forecast.UID0022FullPro
Quadratic regressionProject April 1970 Treasury assets from monthly figures covering October 1969–March 1970.UID0139Full
Cubic regressionFit historical 1989–2013 surplus/deficit figures, project 2025, and compare against a Treasury estimate.UID0140FullPro
Single exponential smoothingForecast 1968 matured noninterest-bearing debt using 1960–1967 data and a specified smoothing factor, α = 0.4.UID0051Full
Compound-growth extrapolationCalculate the annualized compound growth of Series I savings-bond debt between March 2001 and March 2006, then project March 2011 assuming that growth continues.UID0222FullPro
ARIMA(1,1,0)Assemble 1978–1982 bank-held Treasury securities data and apply the specified ARIMA model.UID0198Full

Full without Pro means that particular example is labeled easy and therefore excluded from the hard-only Pro subset. It does not establish that the entire method is absent from Pro. These are examples, not an exhaustive count by algorithm. For revenue specifically, the clearest example is forecasting government tax receipts — not a company's sales pipeline or subscription revenue. The related UID0013, included in Pro, asks for the fitted income-tax regression's slope and intercept rather than the subsequent projection.

3What else the benchmark tests

Forecasting is one strand. The remainder of the question set exercises trend extraction, statistical relationships, risk metrics, external-data adjustment, variance analysis — and the evidence work that precedes any of it.

AreaWhat the questions requireExample
Trend analysis & smoothingHodrick–Prescott filtering to extract trends in federal receipts and outlays, plus moving averages and triangular weighted moving averages. Time-series analysis — but extracting a historical trend is not necessarily forecasting a future value.UID0111 — HP filter with λ = 100 on 2010–2024 receipts and outlays
Statistical relationships & variabilityCorrelation, partial correlation, regression fit measured by R², standard deviation, percentiles, skewness, and kurtosis. This goes beyond simply fitting a trend line.UID0103 — bond-yield relationships, including correlation after controlling for time
Financial risk & growth metricsHistorical Value at Risk, expected shortfall, volatility, growth rates, and weighted averages. These test execution of defined calculations; their sometimes stylized inputs should not be mistaken for validation of a production portfolio-risk system.UID0052 — VaR
UID0069 — historical expected shortfall
Inflation adjustment & currency conversionCombining Treasury figures with external information — a price index or exchange rate — to put amounts on a comparable basis. Tests both sourcing the correct adjustment and applying it correctly, rather than treating every dollar figure as directly comparable.External data must be found and applied to the right base period
Forecast-versus-actual comparisonsSome tasks evaluate an existing projection rather than create one. Closer to budget variance analysis than to time-series model fitting.UID0204 — difference between the projected FY2010 budget deficit and the actual deficit
Finding & reconciling the evidenceLocating the correct tables, reading hierarchical headers, combining documents, distinguishing fiscal from calendar years, and selecting the appropriate revision of a reported number. Databricks documents failures where the method was reasonable but the agent used the wrong periods or source figures.The benchmark's central question: can the agent find the right evidence and perform the requested analysis correctly?

4What a perfect score establishes

The important distinction is between correctly executing a specified forecast calculation and producing a forecast that accurately predicts the real world. Many questions prescribe the method — and sometimes its parameters — and check the final answer against a reference with numerical tolerance.

Established by 100%

  • Document-based quantitative financial analysis — finding the right tables in a large government corpus and reading them correctly.
  • Executing several kinds of forecasts as specified: linear, quadratic and cubic regression; exponential smoothing; compound growth; ARIMA(1,1,0).
  • Statistical reasoning — correlation and partial correlation, R², dispersion and shape statistics, HP-filter trend extraction.
  • Defined risk and growth calculations — VaR, expected shortfall, volatility, weighted averages.
  • Basis adjustment — sourcing and applying price-index and exchange-rate corrections.
  • Variance analysis — comparing projections to realized outcomes.
  • Period and revision discipline — fiscal vs. calendar years, correct data vintage, hierarchical headers.

Not established by this score alone

  • Choosing the best forecasting model for an unfamiliar series — the benchmark prescribes the method.
  • Well-calibrated prediction intervals — answers are point values checked against a reference.
  • Reliably forecasting an unfamiliar company's revenue — the corpus is government receipts, not sales pipelines or subscriptions.
  • An integrated three-statement financial model.
  • A discounted-cash-flow valuation.
  • A revenue forecast built from customer acquisition, pricing, conversion, and churn — these warrant separate tasks and evaluation criteria.
  • A production portfolio-risk system — risk questions use stylized inputs.

How One Jump describes the result

A 100% OfficeQA score is evidence for document-based quantitative financial analysis, including executing several kinds of forecasts. It should not be used alone to claim validation of an integrated three-statement model, a DCF valuation, or a revenue forecast built from acquisition, pricing, conversion, and churn — those warrant separate tasks and evaluation criteria.

Corporate inquiries

Discuss the model with One Jump.

For capability, integration, and enterprise-access questions, contact the company directly.

[email protected]