top of page
Exponential Logo.png

Investor Analytics by Investor Type, Now as a Factor: Introducing XTech Flow Factors

  • Jul 21
  • 4 min read

Two cross-sectional signals from the XTech Flow Factors Library — Institutional Participation Share and Retail Participation Share — built on XTech's participant-classified flow analytics and delivered as pre-normalized factor columns. This note covers what each measures, where it works, and where it does not.


two order flow factors ready to score

For most of its life, XTech's flow analytics rewarded a particular kind of reader: one fluent in order-book microstructure, comfortable turning our raw, minute-level flow feed into a tradable signal and reasoning about how temporary and permanent market impact shape it.


That fluency was the price of entry, and it narrowed the audience to desks with deep experience in market-impact estimation and statistical arbitrage.


XTech Flow Factors Library changes the packaging, not the underlying information content. Our veteran research team has created a full panel of Flow Factors atuned to a full range of investment time horizons and concepts. In this article, we analyze two signals as pre-computed pre-normalized factors columns that drop directly into a standard cross-sectional normalized evaluation — no microstructure reconstruction, no classification pipeline to build.


  • Quant teams score the columns against future returns on day one.

  • Fundamental investors read them as a flow sentiment or market position-level gauge of who is trading a name.


Where the signals come from


Both example factors rest on the same foundation as XTech's flow products: order flow classified by participant type.


Every trade is attributed to institutional or retail participants using microstructure features, and the factor summarizes that flow into a single cross-sectional measure of who is driving a name's volume.


The columns arrive already cross-sectionally rank-gaussianized, so the normalization that usually sits between a raw feed and a usable factor is already done.



The first release is two factors, each isolating one side of that participant split.


  1. Institutional Participation Share reads institutional crowding in small caps

  2. Retail Participation Share reads retail crowding among the most liquid large caps.


They share a foundation but are calibrated to different universes and horizons — the sections below take each in turn, then compare them directly. 


Institutional Participation Share (IPS)


The 20-day exponentially-weighted share of a name's order flow attributed to institutional participants — a crowding measure.


On the Russell 2000, names where flow is institution-dominated outperform retail-dominated names the next trading day, with the short leg contributing most of the spread.


At a glance


What the tests show:


  • Significant at every horizon tested; IC strengthens with horizon (0.017 → 0.024).

  • Same sign in every calendar year since 2022 — no reversals.

  • Confirmed out-of-sample at its native one-day horizon.

  • Not a volume-spike artifact — excluding names trading at 5× average volume moves the spread Sharpe by under 2%.

  • Turnover ~30% daily, ~98% factor-driven (the Russell 2000 reconstitutes annually rather than re-ranking daily).

  • Gross-of-cost Sharpe is 2.40, so net performance is genuinely cost-sensitive at this turnover — hence the conservative 20 bp assumption for a less-liquid small-cap book.



Retail Participation Share (RPS)


The concentration of a name's order flow in retail hands — the mirror image of IPS, and the strongest signal in this flow-factor family to date.


Among the most liquid US names, low-RPS (institution-dominated) names outperform high-RPS, retail-crowded names over the next ten days.


At a glance



What the tests show:


  • Highest Rank ICIR of any factor tested to date.

  • Edge is monotonic across horizons — strengthening from 1 to 5 to 10 days, the signature of a slow-building crowding effect.

  • Confirmed out-of-sample: the holdout matches or exceeds in-sample.

  • Same sign every year since 2022, including the 2022 selloff and 2023 mega-cap rally.

  • Decays gently — Rank IC falls only ~28% by the ten-day rebalance, then holds flat to 30 days, so execution delay costs little.

  • Decile spread: lowest-RPS decile +0.41% vs highest −0.44% over ten days (rank correlation −0.90).

  • On the out-of-sample window alone: 29.1% annualized, 75% hit rate.


The two side by side



The two are complements, not variants of one idea — different universes, horizons, and cadences, pointing in opposite directions on the same question of who is trading a name.


Using the Factors


Systematic desks:


  • Trade each directly as the decile long/short above, at the cadences shown.

  • Or — more likely — drop the pre-normalized columns into an existing factor model as a zero-lag, participant-flow input, valued for its orthogonality to price- and volume-based factors.


Fundamental and discretionary investors:


  • A Russell 2000 name climbing into the top institutional-share decile corroborates an existing long; one sliding to the bottom is a caution flag.

  • Among liquid large caps, a name jumping into the top RPS decile — becoming retail-crowded — flags a long you already hold.

  • Used this way the factors are a confirmation and crowding layer over fundamental theses, not a trading system — the audience the raw feed never served.


Caveats


  • Short out-of-sample window (~8 months). Enough to confirm neither edge decayed, not to pin a precise OOS Sharpe.

  • Commission-only cost model. Borrow, financing, and market impact are not included; a retail-crowded short book can carry a real borrow premium.

  • IPS is cost-sensitive at its ~30% daily turnover — reported at a conservative 20 bp.

  • Capacity only partly characterized — no full liquidity-tier or dollar-capacity study at live allocation size yet.


Statistics use decile bucketing and MAD-winsorization at 3.5; p-values are overlap-robust and Bonferroni-corrected against the full research sweep. None of these caveats say the signals are not real — they are the boundary conditions a research buyer needs to evaluate them.



Comments


Unlock Your Data's Potential Today.

Schedule your free consultation today and discover how we can transform your data strategy.

bottom of page