OfficeQA tests considerably more than ARIMA.
OfficeQA combines document retrieval, financial calculations, statistical analysis, and several forecasting methods across 246 questions answered from a corpus of 89K+ U.S. Treasury Bulletin documents.
Custom neural networks · Financial intelligence
Capability report · Financial-document analysis · Statistical reasoning · Specified-model forecasting
OfficeQA — Full (246 questions) and Pro (133 hard questions)
Corpus — 89K+ U.S. Treasury Bulletin documents
OfficeQA combines document retrieval, financial calculations, statistical analysis, and several forecasting methods across 246 questions answered from a corpus of 89K+ U.S. Treasury Bulletin documents.
Questions specify the model — often its parameters — and check the final number against a reference. Finding the right tables, periods, and revisions is the agent's job, and where most failures happen.
Regression, smoothing, HP filtering, correlation, VaR, expected shortfall, inflation and currency adjustment, and forecast-versus-actual comparisons all appear in the question set.
A perfect score establishes document-based quantitative analysis and specified forecasts. It does not, by itself, establish model selection, calibrated intervals, or company revenue forecasting.
Financial-document analysis, statistical reasoning, and specified-model forecasting — not merely ARIMA forecasting, and not a complete test of autonomous corporate financial planning.
OfficeQA Full and Pro retrieve from the same collection of 89K+ Treasury Bulletin documents.
questions in OfficeQA Full — all answered correctly
hard questions make up OfficeQA Pro — same corpus, easy items excluded
of Pro questions require analysis beyond basic arithmetic (technical paper)
Treasury Bulletin documents to retrieve from — hierarchical tables, fiscal years, revised figures
Concrete questions from the benchmark — not methods an agent could theoretically use. Each prescribes the model, and often its parameters; the answer is checked against a reference value with numerical tolerance.
| Method | What the question asks the agent to do | Question | Appears in |
|---|---|---|---|
| Linear regressionrevenue | Fit federal individual income-tax receipts from 1929–1942, then extrapolate receipts to a later fiscal year. | UID0014 | Full |
| Linear regressionspending | Use Agriculture Department spending from 1990–1998 to project 1999 spending, returning the slope, intercept, and forecast. | UID0022 | FullPro |
| Quadratic regression | Project April 1970 Treasury assets from monthly figures covering October 1969–March 1970. | UID0139 | Full |
| Cubic regression | Fit historical 1989–2013 surplus/deficit figures, project 2025, and compare against a Treasury estimate. | UID0140 | FullPro |
| Single exponential smoothing | Forecast 1968 matured noninterest-bearing debt using 1960–1967 data and a specified smoothing factor, α = 0.4. | UID0051 | Full |
| Compound-growth extrapolation | Calculate the annualized compound growth of Series I savings-bond debt between March 2001 and March 2006, then project March 2011 assuming that growth continues. | UID0222 | FullPro |
| ARIMA(1,1,0) | Assemble 1978–1982 bank-held Treasury securities data and apply the specified ARIMA model. | UID0198 | Full |
Full without Pro means that particular example is labeled easy and therefore excluded from the hard-only Pro subset. It does not establish that the entire method is absent from Pro. These are examples, not an exhaustive count by algorithm. For revenue specifically, the clearest example is forecasting government tax receipts — not a company's sales pipeline or subscription revenue. The related UID0013, included in Pro, asks for the fitted income-tax regression's slope and intercept rather than the subsequent projection.
Forecasting is one strand. The remainder of the question set exercises trend extraction, statistical relationships, risk metrics, external-data adjustment, variance analysis — and the evidence work that precedes any of it.
| Area | What the questions require | Example |
|---|---|---|
| Trend analysis & smoothing | Hodrick–Prescott filtering to extract trends in federal receipts and outlays, plus moving averages and triangular weighted moving averages. Time-series analysis — but extracting a historical trend is not necessarily forecasting a future value. | UID0111 — HP filter with λ = 100 on 2010–2024 receipts and outlays |
| Statistical relationships & variability | Correlation, partial correlation, regression fit measured by R², standard deviation, percentiles, skewness, and kurtosis. This goes beyond simply fitting a trend line. | UID0103 — bond-yield relationships, including correlation after controlling for time |
| Financial risk & growth metrics | Historical Value at Risk, expected shortfall, volatility, growth rates, and weighted averages. These test execution of defined calculations; their sometimes stylized inputs should not be mistaken for validation of a production portfolio-risk system. | UID0052 — VaRUID0069 — historical expected shortfall |
| Inflation adjustment & currency conversion | Combining Treasury figures with external information — a price index or exchange rate — to put amounts on a comparable basis. Tests both sourcing the correct adjustment and applying it correctly, rather than treating every dollar figure as directly comparable. | External data must be found and applied to the right base period |
| Forecast-versus-actual comparisons | Some tasks evaluate an existing projection rather than create one. Closer to budget variance analysis than to time-series model fitting. | UID0204 — difference between the projected FY2010 budget deficit and the actual deficit |
| Finding & reconciling the evidence | Locating the correct tables, reading hierarchical headers, combining documents, distinguishing fiscal from calendar years, and selecting the appropriate revision of a reported number. Databricks documents failures where the method was reasonable but the agent used the wrong periods or source figures. | The benchmark's central question: can the agent find the right evidence and perform the requested analysis correctly? |
The important distinction is between correctly executing a specified forecast calculation and producing a forecast that accurately predicts the real world. Many questions prescribe the method — and sometimes its parameters — and check the final answer against a reference with numerical tolerance.
Established by 100%
Not established by this score alone
A 100% OfficeQA score is evidence for document-based quantitative financial analysis, including executing several kinds of forecasts. It should not be used alone to claim validation of an integrated three-statement model, a DCF valuation, or a revenue forecast built from acquisition, pricing, conversion, and churn — those warrant separate tasks and evaluation criteria.
Corporate inquiries
For capability, integration, and enterprise-access questions, contact the company directly.
[email protected]