Crafting a Trading Strategy

Developing a systematic trading strategy rarely begins with a perfect idea. Instead, it is an iterative process of forming hypotheses, testing them, discarding what does not work, and refining what does. In this article, we develop a futures trading strategy from its underlying economic rationale through successive improvements, illustrating both the research process and the statistical techniques used to evaluate each step.

37 minutes

Alpha rarely arrives as one hero­ic dis­cov­ery. It is more like an ant colony car­ry­ing a leaf many times its own size: dozens of small con­tri­bu­tions, each unim­pres­sive on its own, some­how pro­duc­ing an impres­sive result.

Unfor­tu­nate­ly, research also resem­bles an ant colony in anoth­er respect. Much of the work involves run­ning in cir­cles.

The first ver­sion of a sys­tem­at­ic strat­e­gy almost nev­er sur­vives unchanged. Every iter­a­tion expos­es anoth­er assump­tion, anoth­er weak­ness, or anoth­er oppor­tu­ni­ty for a small improve­ment. Good strate­gies are not dis­cov­ered ful­ly formed. They are craft­ed.

To har­vest such incre­men­tal improve­ments, you need a dis­ci­plined research process. Fol­low­ing it step by step helps avoid the dan­gers of over­fit­ting.

Before devel­op­ing the strat­e­gy, let’s intro­duce the under­ly­ing research process.

Research Process

Fig­ure 1 illus­trates the research process. Notice that every arrow in Fig­ure 1 even­tu­al­ly points back to an ear­li­er stage. Research is rarely lin­ear. Every improve­ment uncov­ers anoth­er ques­tion.

Research process for a trad­ing strat­e­gy

The process is cycli­cal because every stage may lead to new con­clu­sions. If this is the case, any step may require return­ing to an ear­li­er one.

The research process con­sists of sev­en stages and is influ­enced by First Prin­ci­ples Think­ing.

Economic Rationale

To min­i­mize the risk of over­fit­ting, do not start by crunch­ing data and tun­ing para­me­ters. Start with a sound eco­nom­ic ratio­nale. Write the ratio­nale down. Make every assump­tion explic­it. You will almost cer­tain­ly revis­it both lat­er.

Every sub­se­quent mod­i­fi­ca­tion should be jus­ti­fied by the eco­nom­ic ratio­nale. If it does not, dis­card it. Of course, not every eco­nom­ic ratio­nale will prove cor­rect or be exploitable in a sys­tem­at­ic way. Keep the ratio­nale in mind through­out the process. Only make changes to a strat­e­gy that are in line with the eco­nom­ic ratio­nale.

To illus­trate this phase, let’s turn to a prac­ti­cal strat­e­gy: We want to par­tic­i­pate in trend­ing assets and for­mu­late an eco­nom­ic ratio­nale.

Eco­nom­ic Ratio­nale

Finan­cial mar­kets con­tin­u­ous­ly process new infor­ma­tion about the econ­o­my. While news may arrive instant­ly, its impact on prices often unfolds grad­u­al­ly. Mar­ket par­tic­i­pants dif­fer in their infor­ma­tion, objec­tives, risk con­straints, and speed of deci­sion-mak­ing. As a result, cap­i­tal is real­lo­cat­ed over time rather than all at once. This grad­ual adjust­ment cre­ates per­sis­tent price move­ments that can last for weeks, months, or even years. Trend-fol­low­ing does not attempt to pre­dict these move­ments before they begin. Instead, it assumes that once a per­sis­tent adjust­ment is under­way, it is more like­ly to con­tin­ue than to reverse imme­di­ate­ly.

Signal Design

A sound eco­nom­ic ratio­nale is necessary—but it can­not be trad­ed. The next step is to trans­late the idea into a sig­nal that can be eval­u­at­ed objec­tive­ly.

A sig­nal is a for­mu­la that assigns a numer­i­cal val­ue to each instru­ment1 at each point in time. As we are devel­op­ing a trend strat­e­gy for futures on dai­ly price data, we need a sig­nal for every trad­ing day. Oth­er strate­gies may instead rely on Com­mit­ments of Traders (COT) data, intra­day data, or alter­na­tive data sources.

When defin­ing a sig­nal, it is impor­tant to think about poten­tial look-ahead bias­es. Is every part of the sig­nal data avail­able at every point in time? The sig­nal must not use infor­ma­tion about the future. Look-ahead bias often creeps in through sur­pris­ing­ly sub­tle mis­takes.

For exam­ple, you may devise a sig­nal that uses GDP data. As GDP data often gets revised after its first pub­li­ca­tion it is cru­cial to only use the data that was actu­al­ly avail­able at the time the sig­nal is cre­at­ed. Anoth­er com­mon mis­take is to fill miss­ing data by prop­a­gat­ing future data back­ward.

Nev­er do this. Seri­ous­ly. Always prop­a­gate miss­ing data from the last known val­ue for­ward. If there is miss­ing data at the begin­ning of a time series and for­ward prop­a­gat­ing is not pos­si­ble, then set the sig­nal to a neu­tral val­ue until the first observed val­ue. This pop­u­lar mis­take reg­u­lar­ly occurs when cal­cu­lat­ing stan­dard devi­a­tions with a giv­en look­back peri­od.

The val­ue of the sig­nal indi­cates the con­vic­tion in a trad­ing posi­tion: You want to be long for high val­ues (per­haps even in a larg­er posi­tion size) and short for low val­ues.

The sig­nal should be con­struct­ed so that it is com­pa­ra­ble across your trad­able uni­verse. For exam­ple, in the case of a trend strat­e­gy, Bit­coin should not have a high­er sig­nal val­ue only because it is more volatile than bonds.

When defin­ing a sig­nal it is impor­tant to keep in mind that rapid­ly chang­ing sig­nal val­ues will result in high port­fo­lio turnover: If a sig­nal switch­es its val­ue between long and short every day, it might be a very good sig­nal but the strat­e­gy becomes imprac­ti­cal to trade due to high trans­ac­tion costs because of dai­ly rebal­anc­ing.

The sig­nal def­i­n­i­tion should also spec­i­fy a pos­si­ble range of para­me­ters to ana­lyze. This range may depend on the eco­nom­ic ratio­nale or avail­able data.

For exam­ple, if you want to cap­ture trends there might be an upper bound for the trend length because of eco­nom­ic rea­sons or sim­ply the length of his­tor­i­cal data: It does not make sense to ana­lyze 10-year trends if you only have 5 years of his­tor­i­cal data.

To cre­ate a sig­nal for our trend-fol­low­ing strat­e­gy, we need to iden­ti­fy trends in prices. Math­e­mat­i­cal­ly, there are count­less ways to iden­ti­fy a trend in a time series: You may fit a regres­sion on the log prices, cre­ate a time series mod­el, con­sult some tech­ni­cal indi­ca­tors like the rel­a­tive strength index or do mov­ing-aver­age crossovers.

For rea­sons of sim­plic­i­ty, we will stick with the well-known mov­ing aver­ages. To be pre­cise, we will use a spe­cial vari­ant called expo­nen­tial­ly weight­ed mov­ing aver­age (EWMA). This vari­ant gives more recent obser­va­tions a greater weight so that the sig­nal will respond more quick­ly to new infor­ma­tion.

Sig­nal Design

Cal­cu­late a trad­ing sig­nal for each trad­able instru­ment with the fol­low­ing for­mu­la:


\(\mathsf{signal_t = \displaystyle\frac{trend_t(P_t)}{{\sigma_{t}}}\times Vol_t}\)

\(\mathsf{P_t}\): Back-adjust­ed price series of an asset at time t

\(\mathsf{\sigma_t}\): Price volatil­i­ty of the asset at time t

\(\mathsf{trend_t(P_t)}\) is a func­tion to cal­cu­late the trend strength of an asset at time t:

\[\mathsf{trend_t(P_t) = ewma_{fast}(P_t)-ewma_{slow}(P_t)}\]

\(\mathsf{ewma_{span}(P_t)}\): The expo­nen­tial mov­ing-aver­age price of the back-adjust­ed price time series \(\mathsf{P_t}\) with a look­back peri­od of span trad­ing days. The look­back is spec­i­fied as a span para­me­ter of pan­das’ ewm method. The span in Python is defined in a way to make it com­pa­ra­ble to the same look­back of a sim­ple mov­ing aver­age. The decay can equiv­a­lent­ly be expressed using the smooth­ing fac­tor \(\alpha\) or the half-life; these para­me­ter­i­za­tions can be con­vert­ed into one anoth­er. We always set the slow span to four times the fast span to reduce the num­ber of free para­me­ters to choose from.

\(\mathsf{Vol_t}\): Volatil­i­ty mul­ti­pli­er of an asset at time t. Divi­sion of the long-term volatil­i­ty (smoothed short-term volatil­i­ty over 20 years) by the cur­rent short-term volatil­i­ty (price volatil­i­ty of the asset smoothed with a span of 32 days). The val­ue is then per­centile-mapped to a range between 0.5 and 2.

Let’s decom­pose the for­mu­las on an abstract lev­el: The trad­ing sig­nal nor­mal­izes the trend strength of an asset’s price \(\mathsf{trend_t(P_t)}\) at time t by its price volatil­i­ty \(\mathsf{\sigma_t}\) and is then scaled by a time-depen­dent fac­tor \(\mathsf{Vol_t}\).

Divid­ing the trend strength by the volatil­i­ty is called inverse volatil­i­ty scal­ing and serves sev­er­al pur­pos­es: As the trend func­tion takes back-adjust­ed prices as input the trend strength will be depen­dent on the mag­ni­tude of an absolute price move­ment. These move­ments depend large­ly on the volatil­i­ty of an asset. As the sig­nal should be com­pa­ra­ble across assets of dif­fer­ent asset class­es, this nor­mal­izes the sig­nal across instru­ments.

By divid­ing by the volatil­i­ty, an asset like Bit­coin (which is very volatile) can be com­pared to an asset like the 10-year U.S. bonds.

Addi­tion­al­ly, inverse-volatil­i­ty scal­ing has anoth­er advan­tage: The sig­nal decreas­es when trend strength remains unchanged but volatil­i­ty ris­es. This way we also get a risk man­age­ment mech­a­nism built-in for free! As the sig­nal is trans­lat­ed into a posi­tion size of our port­fo­lio, the posi­tion size will decrease when the same trend becomes riski­er. A wel­come bonus.

Next, con­sid­er the trend-strength func­tion. \(\mathsf{trend_t}\): It is the dif­fer­ence of two mov­ing aver­ages fast and slow. These are two para­me­ters of our sig­nal. The fast com­po­nent mea­sures the mov­ing-aver­age of a price series with a short­er look­back than the slow com­po­nent. Sup­pose an asset appre­ci­at­ed sub­stan­tial­ly over the pre­vi­ous 10 trad­ing days. Set the fast span to 10 days and the slow span to 40 days. If the asset did not move much the pre­ced­ing 30 trad­ing days, then \(\mathsf{ewma_{10}}\) will be high­er than \(\mathsf{ewma_{40}}\).

\(\mathsf{trend_t}\) will there­fore have a pos­i­tive val­ue and sig­nal a long posi­tion. The trend func­tion can be inter­pret­ed eco­nom­i­cal­ly as the short-term price lev­el com­pared to a longer-term equi­lib­ri­um price. Pos­i­tive val­ues sig­nal price appre­ci­a­tion and neg­a­tive val­ues price declines com­pared to the equi­lib­ri­um price. This inter­pre­ta­tion direct­ly reflects our eco­nom­ic ratio­nale: the mar­ket is still adjust­ing toward a new equi­lib­ri­um.

We also need to spec­i­fy a prac­ti­cal range of the look­back peri­od. This deci­sion is less dri­ven by the eco­nom­ic ratio­nale but more by the avail­able data. Because the trad­able uni­verse expands over time, it increas­ing­ly includes instru­ments with­out data his­to­ries extend­ing back to the 1970s.

To pre­serve diver­si­fi­ca­tion, we do not want to wait until every instru­ment has accu­mu­lat­ed sev­er­al years of obser­va­tions. We there­fore cap the fast EWMA span at 150 trad­ing days, cor­re­spond­ing to a slow span of 600 trad­ing days, or approx­i­mate­ly 2.4 years. We set the low­er bound of the fast span at two trad­ing days because faster sig­nals would gen­er­al­ly pro­duce exces­sive turnover and trans­ac­tion costs.

What about the volatil­i­ty mul­ti­pli­er \(\mathsf{Vol_t}\)?

Volatil­i­ty Mul­ti­pli­er of the U.S. 10-year Bond future between 2010 and 2018. Cal­cu­lat­ed as the expo­nen­tial­ly weight­ed short-term volatil­i­ty with a span of 20 years divid­ed by the same quan­ti­ty with a span of 32 trad­ing days. If there is less than 20 years of data avail­able the max­i­mum avail­able span will be used every trad­ing day.

This scal­ing fac­tor adjusts the sig­nal strength on a long-term lev­el. Con­sid­er the fol­low­ing sit­u­a­tion in the after­math of the finan­cial cri­sis in Fig­ure 2: Inter­est rates were arti­fi­cial­ly sup­pressed by cen­tral banks for the bet­ter part of a decade.

Inter­est volatil­i­ty was also quite low dur­ing that time. The volatil­i­ty mul­ti­pli­er there­fore remained above 1 for much of this peri­od. The trad­ing sig­nal will there­fore be high­er than nor­mal and increase the posi­tion size of the asset in the port­fo­lio to reflect this unusu­al­ly low-volatil­i­ty envi­ron­ment.

Signal Analysis

At this stage, we delib­er­ate­ly avoid look­ing at back­test results. Instead, ver­i­fy how the cal­cu­lat­ed sig­nals behave across the trad­able uni­verse.

Are there out­liers due to data errors or spe­cial cir­cum­stances? What is the gen­er­al range of the sig­nals over all instru­ments? Do they behave sim­i­lar­ly across dif­fer­ent asset class­es? How smooth are the trad­ing sig­nals from day to day?

Many of these checks con­sist of visu­al inspec­tion of the sig­nals of dif­fer­ent instru­ments but some aspects can also be checked quan­ti­ta­tive­ly. Any unex­pect­ed behav­ior should be inves­ti­gat­ed and cor­rect­ed before pro­ceed­ing to the next step.

Let’s dive direct­ly into the trend-fol­low­ing exam­ple and first have a look at the num­ber of trad­able instru­ments over time in Fig­ure 3. The his­tor­i­cal data of the uni­verse starts in June 1975. Back then there were only a few estab­lished con­tracts such as soy­bean and U.S. Trea­sury futures.

Over time, the futures mar­ket became increas­ing­ly diverse. New instru­ments entered the mar­ket; some suc­ceed­ed, while oth­ers, such as pork-bel­ly futures, were dis­con­tin­ued. In aggre­gate, more and more instru­ments have been trad­able over time.

The con­tin­ued expan­sion of the uni­verse con­strains the max­i­mum look­back we can eval­u­ate as we want to trade as many instru­ments as pos­si­ble. Under our “burn-in” rule, Trend 150 requires 600 trad­ing days of data before it becomes active. The 600-day cut­off rep­re­sents a com­pro­mise between requir­ing suf­fi­cient his­to­ry and retain­ing as broad a trad­able uni­verse as pos­si­ble.

Num­ber of trade­able instru­ments over time. An instru­ment is clas­si­fied as trade­able if it sat­is­fies mar­ket data and liq­uid­i­ty con­straints. Smoothed for read­abil­i­ty.

To make the trad­ing sig­nal more con­crete, con­sid­er Fig­ure 4. The fig­ure depicts trend sig­nals with dif­fer­ent look­back peri­ods. Lean hogs, for exam­ple, had a high val­ue in June 2014 on the long-term Trend 110. That means the short­er-term EWMA was sub­stan­tial­ly above the longer-term EWMA, indi­cat­ing strong recent appre­ci­a­tion rel­a­tive to the longer-term price lev­el.

One year lat­er the trend reversed over that hori­zon, putting the trad­ing sig­nal in neg­a­tive ter­ri­to­ry in June 2015. The three exam­ples show­case anoth­er fea­ture of dif­fer­ent trend look­backs: The longer the look­back, the more slow­ly the whole sig­nal changes. This also implies low­er trad­ing costs when actu­al­ly trad­ing the sig­nal.

To get a bet­ter feel­ing for all trad­ing sig­nals over all instru­ments we will now turn to Fig­ure 5 and Fig­ure 6. These fig­ures show dif­fer­ent val­ues of the trad­ing sig­nal per look­back peri­od over all instru­ments.

Fig­ure 5 shows the mean absolute sig­nal, the 5th–95th per­centile band, and the full minimum–maximum range. You may ask your­self why all val­ues increase as the look­back length­ens.

Trend sig­nals with fast-EWMA spans of 30, 80, and 110 trad­ing days for U.S. 10-year Trea­sury futures, the S&P 500 Index, and lean hog futures. The cor­re­spond­ing slow-EWMA spans are 120, 320, and 440 trad­ing days. Sig­nals are named after their fast span.

That is by con­struc­tion of the sig­nal. The volatil­i­ty esti­mate is held con­stant across look­backs rather than scaled to the sig­nal hori­zon. As the look­back length­ens, the poten­tial mag­ni­tude of the trend com­po­nent increas­es while the volatil­i­ty denom­i­na­tor remains unchanged. The result­ing sig­nal val­ues there­fore tend to increase.

This effect could be removed by scal­ing the volatil­i­ty more to the length of the look­back peri­od (mul­ti­ply by \(\mathsf{\sqrt{(slow-fast)/2}}\)).

But what actu­al­ly mat­ters much more is the smooth­ness of the mean, per­centile, and minimum–maximum bands: In the fig­ure there are no spikes.

Spikes may indi­cate a flaw in the sig­nal design, a data prob­lem, or an insuf­fi­cient ini­tial­iza­tion peri­od. Ini­tial­iza­tion (“burn-in”) means that the sig­nal is only cal­cu­lat­ed after a com­plete look­back win­dow is avail­able. Depend­ing on how your sig­nal is designed, a miss­ing “burn-in” may result in extreme val­ues at the begin­ning of the time series.

Val­ues of the (absolute) trad­ing sig­nal of the trend strat­e­gy with dif­fer­ent look­backs over the trad­able uni­verse of instru­ments. Neu­tral sig­nal val­ues due to ”burn-in” of a time series have been removed.

Fig­ure 6 is the first fig­ure that can real­ly embar­rass us. If one asset class sud­den­ly pro­duces sig­nal val­ues twice as large as every­thing else, some­thing is prob­a­bly wrong—either with the data or with our assump­tions. Two asset class­es are worth­while to dis­cuss here: Volatil­i­ty and Rates.

Both asset class­es exhib­it more extreme mean absolute sig­nal val­ues than the oth­ers. Why?

Rates are rel­a­tive­ly straight­for­ward to explain. Our sam­ple begins in 1975 and includes a mul­ti-decade decline in yields. Because many rate-futures prices move inverse­ly to yields, these con­tracts expe­ri­enced per­sis­tent price trends that pro­duced com­par­a­tive­ly strong trend sig­nals.

Volatil­i­ty requires a dif­fer­ent expla­na­tion. First, the asset class con­tains only two instruments—VIX futures in the Unit­ed States and VSTOXX futures in Europe—so its cross-sec­tion­al mean is less sta­ble. Sec­ond, I exclude front-month volatil­i­ty futures because of their ele­vat­ed risk and instead trade more deferred matu­ri­ties.

Out­side peri­ods of mar­ket stress, volatil­i­ty-futures curves are fre­quent­ly in con­tan­go (upward slop­ing). When this occurs, deferred con­tracts tend to decline as they approach expiry and con­verge toward low­er near­by prices, pro­duc­ing neg­a­tive roll yield. This struc­tur­al drift can cre­ate a stronger trend com­po­nent in a con­tin­u­ous futures series.

Mean absolute trad­ing sig­nal by asset class with dif­fer­ent look­backs. Neu­tral sig­nal val­ues due to “burn-in” of a time series have been removed.

Parameter Exploration

The sig­nal looks plau­si­ble, but we still have no idea whether it actu­al­ly trades well. Only now do we allow our­selves to run the first back­test. But we want to glean as lit­tle from the data as pos­si­ble and not look at full-blown equi­ty curves or draw­down plots.

In gen­er­al, it is good prac­tice to run as few back­tests as pos­si­ble. A strat­e­gy rarely works right out of the gate and needs some adjust­ments, but you should nev­er get caught in a process of fid­dling with dif­fer­ent sig­nals.

Remem­ber that if you test hun­dreds of dif­fer­ent sig­nal def­i­n­i­tions and end up with one good back­test, you may have cre­at­ed some­thing that fits your data and gen­er­al­izes very poor­ly.

Exam­in­ing the para­me­ter sur­face helps to avoid this kind of opti­miza­tion trap. Run back­tests for all para­me­ter val­ues and plot the Sharpe ratio (before and after costs) of each result­ing back­test. Per­for­mance should vary smooth­ly across neigh­bor­ing para­me­ter val­ues rather than jump ran­dom­ly.

The goal is not to find one per­fect para­me­ter. It is to com­bine many good ones that rein­force each oth­er for more sta­bil­i­ty, robust­ness and some added diver­si­fi­ca­tion. Like an ant colony, we are not look­ing for one hero­ic work­er but for many that each con­tribute a lit­tle.

This rather the­o­ret­i­cal dis­cus­sion can be made more acces­si­ble by look­ing at the exam­ple of the trend-fol­low­ing strat­e­gy in Fig­ure 7.

Sharpe ratio of the trend-fol­low­ing strat­e­gy with dif­fer­ent look­back peri­ods. The look­back ranges from 2 to 150 for the fast EWMA. The Sharpe ratio is cal­cu­lat­ed with­out a risk-free rate. Costs include com­mis­sions, spreads and the cost of roll-exe­cu­tion.

Fig­ure 7 shows the strategy’s Sharpe ratio before and after esti­mat­ed costs for each fast-EWMA span. As expect­ed, costs reduce per­for­mance across the para­me­ter range. The before- and after-cost curves dif­fer by approx­i­mate­ly 0.15 Sharpe-ratio points across most look­backs, reflect­ing com­mis­sions, bid–ask spreads, and roll-exe­cu­tion costs. At the short­est look­backs, high turnover mate­ri­al­ly depress­es after-cost per­for­mance. This effect dimin­ish­es once the fast span reach­es rough­ly 10 trad­ing days.

The his­tor­i­cal Sharpe ratio then declines grad­u­al­ly as the look­back length­ens. This does not nec­es­sar­i­ly make slow­er sig­nals unat­trac­tive: they trade less fre­quent­ly and may diver­si­fy faster spec­i­fi­ca­tions. More impor­tant than max­i­miz­ing per­for­mance at a sin­gle para­me­ter val­ue is the smooth shape of the curve. There is no iso­lat­ed opti­mum; neigh­bor­ing look­backs pro­duce sim­i­lar results, which sug­gests that per­for­mance is not dri­ven by one nar­row­ly select­ed para­me­ter.

Cor­re­la­tions between returns of the trend-fol­low­ing strat­e­gy with look­back peri­ods for EWMA fast from 2 to 150 (step size varies).

Fig­ure 8 shows the cor­re­la­tions between strat­e­gy returns gen­er­at­ed by dif­fer­ent fast-EWMA spans. Cor­re­la­tions change rapid­ly among the short­est look­backs but approach 100% among the longest. For exam­ple, Trend 2 and Trend 10 have a cor­re­la­tion of approx­i­mate­ly 46%, where­as Trend 100 and Trend 120 have a cor­re­la­tion of approx­i­mate­ly 99%. Longer EWMAs have strong­ly over­lap­ping weight­ing ker­nels and there­fore respond sim­i­lar­ly to price move­ments. This sup­ports using a denser para­me­ter grid at short hori­zons and a coars­er grid at long hori­zons, where addi­tion­al spec­i­fi­ca­tions pro­vide lit­tle incre­men­tal diver­si­fi­ca­tion.

Parameter Selection

This step in the frame­work deter­mines the range of para­me­ters to trade with the strat­e­gy. A visu­al inspec­tion of the para­me­ter sur­face helps iden­ti­fy robust ranges. With­in this range, an algo­rithm can select the para­me­ters to trade. The whole process is explained in detail in this arti­cle.

This will be the first time we actu­al­ly look at an equi­ty curve in the whole work­flow. But remem­ber: A promis­ing back­test still proves almost noth­ing.

The algo­rithm to select the look­backs to trade sug­gests the com­bi­na­tion of look­backs 2, 4, 6, 10, 20, 35, 60 and 100 trad­ing days (EWMA fast) with a max­i­mum cor­re­la­tion of 95% between adja­cent select­ed look­backs. The algo­rithm was instruct­ed to search a com­bi­na­tion of look­backs between 2 and 100 trad­ing days. The low­er bound was select­ed for added ben­e­fit of diver­si­fi­ca­tion while the upper bound was select­ed for a min­i­mum Sharpe ratio of 1.0.

Fig­ure 9 shows the return cor­re­la­tions for the select­ed look­backs. Adja­cent select­ed look­backs are strong­ly cor­re­lat­ed but the cor­re­la­tions drop off rapid­ly the far­ther apart two look­backs are. The algo­rithm also selects more strate­gies with short look­backs than longer ones. That is pre­cise­ly the rea­son why many EWMA-style sys­tems use expo­nen­tial­ly spaced look­backs like EWMA 4, 8, 16, 32 and 64.

Cor­re­la­tions between strat­e­gy returns with a mix of look­back peri­ods (EWMA fast).

Until now, we have delib­er­ate­ly avoid­ed the final equi­ty curve. This is not an act of extra­or­di­nary self-con­trol. It is an attempt to pre­vent the back­test from influ­enc­ing deci­sions that should be based on the signal’s eco­nom­ic ratio­nale and behav­ior.

The sig­nal looks plau­si­ble. The para­me­ters behave smooth­ly. No sin­gle look­back is excep­tion­al. That is pre­cise­ly the point. Good colonies do not rely on a sin­gle very strong ant.

Now the strat­e­gy earns the right to dis­ap­point us.

Fig­ure 10 con­tains sev­er­al impor­tant obser­va­tions. The diver­si­fied trend-fol­low­ing strat­e­gy per­forms strong­ly over the full sam­ple, par­tic­u­lar­ly from the begin­ning of the back­test through the 2008 finan­cial cri­sis. Per­for­mance weak­ens notice­ably there­after.

Sev­er­al CTA man­agers have sug­gest­ed that the low-inter­est-rate envi­ron­ment from 2010 to 2020 con­tributed to the weak­er results.

How­ev­er, this expla­na­tion does not ful­ly hold up in hind­sight. The post-pan­dem­ic infla­tion shock gen­er­at­ed strong trend-fol­low­ing returns, while per­for­mance weak­ened again after 2022 with inter­est rates stay­ing high.

Equi­ty curve and Draw­down of trend-fol­low­ing strat­e­gy with a com­bi­na­tion of look­back peri­ods of 2, 4, 6, 10, 20, 35, 60 and 100 trad­ing days (EWMA fast). No com­pound­ing, start­ing cap­i­tal is 1 and risk tar­get 25%.

Oth­er pos­si­ble expla­na­tions include increased com­pe­ti­tion, changes in mar­ket struc­ture, and ordi­nary vari­a­tion in the returns of a strat­e­gy with an unsta­ble expect­ed return. The evi­dence pre­sent­ed here can­not deter­mine which expla­na­tion, if any, is cor­rect.

Tables 1 and 2 quan­ti­fy the dete­ri­o­ra­tion. The annu­al­ized return falls from 33.39% over the full sam­ple to 14.72% after 2008, while the Sharpe ratio declines from 1.37 to 0.65. The max­i­mum draw­down and max­i­mum draw­down dura­tion are unchanged because the worst episode occurs with­in the post-2008 sam­ple. Month­ly skew­ness improves mod­est­ly.

Sta­tis­tic Val­ue
Return after Costs [% p.a.] 33.39
Sharpe 1.37
Max­i­mum Draw­down [%] 66.03
Avg Draw­down Dura­tion [days] 26.12
Max Draw­down Dura­tion [days] 1066
Month­ly Skew 0.37
Low­er Tail 1.83
Upper Tail 1.59
Win­ning Days [%] 55.05
Per­for­mance sta­tis­tics for the trend-fol­low­ing strat­e­gy using the select­ed mix of look­backs, July 1975–July 2026
Sta­tis­tic Val­ue
Return after Costs [% p.a.] 14.72
Sharpe 0.65
Max­i­mum Draw­down [%] 66.03
Avg Draw­down Dura­tion [days] 51.65
Max Draw­down Dura­tion [days] 1066
Month­ly Skew 0.50
Low­er Tail 2.07
Upper Tail 1.61
Win­ning Days [%] 54.32
Per­for­mance sta­tis­tics for the trend-fol­low­ing strat­e­gy using the select­ed mix of look­backs, Jan­u­ary 2008–July 2026

These results are sober­ing, but they do not estab­lish that clas­si­cal trend fol­low­ing is dead. The strong returns dur­ing the post-pan­dem­ic infla­tion shock show that the strat­e­gy can still ben­e­fit from per­sis­tent macro­eco­nom­ic trends. Nev­er­the­less, an after-cost Sharpe ratio of 0.65 may be insuf­fi­cient for a stand-alone strat­e­gy once esti­ma­tion uncer­tain­ty and imple­men­ta­tion risk are con­sid­ered.

Fig­ure 10 also illus­trates a famil­iar les­son that you may have read a thou­sand times: Past per­for­mance is not indica­tive of future results, espe­cial­ly if you invest­ed in this strat­e­gy before 2022.

I see the draw­down after 2022 as an indi­ca­tor of that. The recov­ery since May 2025 is encour­ag­ing, although it is too ear­ly to deter­mine whether it rep­re­sents a last­ing improve­ment.

This base­line mod­el has not failed spec­tac­u­lar­ly. That would almost be eas­i­er. Instead, it has failed polite­ly: returns remain pos­i­tive, the eco­nom­ic ratio­nale remains plau­si­ble, and the per­for­mance is just mediocre enough to be annoy­ing.

We there­fore return to the research loop and ask a nar­row­er ques­tion: where has the dete­ri­o­ra­tion occurred?

Finding an Alternative

To get a bet­ter han­dle on the prob­lem, it is advis­able to try to find a rea­son for the declin­ing per­for­mance.

As can be seen from Fig­ure 10 the prob­lem can­not be attrib­uted to a spe­cif­ic asset class. It also is not asso­ci­at­ed with spe­cif­ic futures con­tracts, like new­er con­tracts.

The quants over at Quan­ti­ca came up with an inter­est­ing analy­sis in a recent arti­cle: They con­clud­ed that the per­for­mance of trend-sys­tems degrad­ed some­what for short­er look­backs begin­ning 2023.

This obser­va­tion holds true for our strat­e­gy and trad­able uni­verse as can be seen in Fig­ure 11.

How can this sit­u­a­tion be inter­pret­ed eco­nom­i­cal­ly? The eco­nom­ic envi­ron­ment of 2023 and lat­er was espe­cial­ly influ­enced by uncer­tain­ty in the ener­gy mar­kets because of wars (Ukraine and Iran), chang­ing tar­iffs between large eco­nom­ic blocs (Unit­ed States, Europe and Chi­na) and hopes of pro­duc­tiv­i­ty mir­a­cles due to the AI boom.

Equi­ty curve of trend-fol­low­ing strate­gies for short­er and longer look­back peri­ods of 10, 60 and 100 trad­ing days (EWMA fast). Jan 2023 through July 2026, no com­pound­ing, start­ing cap­i­tal is 1 and risk tar­get 25%.

All these fac­tors result in high­er volatil­i­ty and reg­u­lar repric­ing of assets.

High volatil­i­ty may indi­cate that the mar­ket is pro­cess­ing infor­ma­tion and mov­ing between price regimes more rapid­ly. In such an envi­ron­ment, the fixed slow EWMA can remain anchored to obser­va­tions from an increas­ing­ly irrel­e­vant regime. The result­ing trend sig­nal then com­pares the cur­rent price with a stale ref­er­ence lev­el and may remain active long after the orig­i­nal repric­ing has end­ed.

Short­en­ing the slow EWMA when volatil­i­ty ris­es allows the ref­er­ence price to adapt more rapid­ly. This caus­es sig­nals gen­er­at­ed by iso­lat­ed shocks to decay soon­er, while per­sis­tent price move­ments con­tin­ue to pro­duce direc­tion­al expo­sure.

This hypoth­e­sis sug­gests that the orig­i­nal eco­nom­ic ratio­nale is incom­plete. It assumes not only that price adjust­ments per­sist, but also that his­tor­i­cal price obser­va­tions remain rel­e­vant for a con­stant amount of time. If the speed of mar­ket adjust­ment varies with volatil­i­ty, the ref­er­ence hori­zon of the sig­nal should vary as well.

Exten­sion of the Eco­nom­ic Ratio­nale

Finan­cial mar­kets do not process infor­ma­tion at a con­stant speed. Dur­ing sta­ble peri­ods, eco­nom­ic con­di­tions and expec­ta­tions often change grad­u­al­ly, allow­ing his­tor­i­cal prices to remain rel­e­vant for rel­a­tive­ly long peri­ods.

Dur­ing peri­ods of ele­vat­ed uncer­tain­ty, expec­ta­tions are revised more fre­quent­ly and prices are repriced more rapid­ly. Old­er obser­va­tions may then reflect eco­nom­ic assump­tions that are no longer rep­re­sen­ta­tive of cur­rent con­di­tions.

Volatil­i­ty is there­fore inter­pret­ed as an observ­able indi­ca­tion of the pace of mar­ket adjust­ment. When volatil­i­ty is ele­vat­ed, the use­ful life­time of his­tor­i­cal price infor­ma­tion may be short­er; when con­di­tions are sta­ble, old­er prices may remain infor­ma­tive for longer.

This ratio­nale leads direct­ly to a mod­i­fied sig­nal design. The fast EWMA remains respon­si­ble for mea­sur­ing the recent price lev­el, while the span of the slow EWMA is short­ened when cur­rent volatil­i­ty is high rel­a­tive to its longer-term lev­el. The ref­er­ence price there­fore adapts more quick­ly pre­cise­ly when his­tor­i­cal obser­va­tions are assumed to lose rel­e­vance more rapid­ly.

Updat­ed Sig­nal Design

Update the trend com­po­nent of the orig­i­nal for­mu­la with this spec­i­fi­ca­tion:

\[\mathsf{trend_t = ewma_{fast}(P_t)-ewma_{adaptive}(P_t)}\]

\[
\mathsf{adaptive = min(1.0, Vol_t)\times slow}
\]

\(\mathsf{adaptive}\) is a dynam­i­cal­ly adjust­ed slow span depend­ing on the cur­rent volatil­i­ty of each asset. Because \(\mathsf{Vol_t}\) is mapped to the range 0.5–2.0, the adap­tive slow span can fall to half its base­line val­ue when cur­rent volatil­i­ty is unusu­al­ly high.

This rule intro­duces no inde­pen­dent­ly opti­mized para­me­ter. It reuses the exist­ing volatil­i­ty mul­ti­pli­er as an indi­ca­tor of the cur­rent volatil­i­ty regime. Addi­tion­al para­me­ters would increase the risk of over­fit­ting and should be avoid­ed unless they have a clear eco­nom­ic jus­ti­fi­ca­tion.

With an updat­ed eco­nom­ic ratio­nale and an updat­ed sig­nal, the research process con­tin­ues at step 3, where the result­ing sig­nal is ana­lyzed for any anom­alies.

The first result is vis­i­ble in Fig­ure 12 before we run anoth­er back­test: the sig­nal dis­tri­b­u­tion con­tracts com­pared to the base­line mod­el in Fig­ure 5, par­tic­u­lar­ly at longer look­backs. Why? When volatil­i­ty short­ens the slow EWMA, the fast and slow ref­er­ence prices can­not drift as far apart.

Val­ues of the (absolute) trad­ing sig­nal of the adap­tive trend strat­e­gy with dif­fer­ent look­backs over the trad­able uni­verse of instru­ments. Neu­tral sig­nal val­ues due to ”burn-in” of a time series have been removed.

The sig­nal plot also shows no spikes or any oth­er abnor­mal behav­ior, so that we can resume the next stage of the research process: para­me­ter explo­ration.

Sharpe ratio of the adap­tive trend-fol­low­ing strat­e­gy vs the base­line trend-fol­low­ing strat­e­gy with dif­fer­ent look­back peri­ods. The look­back ranges from 2 to 150 for the fast EWMA. The Sharpe ratio is cal­cu­lat­ed with­out a risk-free rate. Costs include com­mis­sions, spreads and the cost of roll-exe­cu­tion.

The result is shown in Fig­ure 13. The adap­tive ver­sion of trend-fol­low­ing improves the after-cost Sharpe ratio by approx­i­mate­ly 0.05 points across most look­backs.

The result is encour­ag­ing. The improve­ment is small but remark­ably con­sis­tent across the para­me­ter range. The adap­tive strategy’s Sharpe-ratio pro­file is smooth and runs almost par­al­lel to the base­line pro­file.

We can there­fore pro­ceed direct­ly to para­me­ter selec­tion.

The selec­tion algo­rithm adds the look­back 3 in addi­tion to the look­backs of the base­line mod­el. In total, the fol­low­ing look­backs are select­ed: 2, 3, 4, 6, 10, 20, 35, 60 and 100 trad­ing days for EWMA fast. Remem­ber that Trend 100 now uses an adap­tive look­back for EWMA slow. Its default val­ue is 400 and can go down to 200 trad­ing days when the volatil­i­ty is very high.

Equi­ty curves for the adap­tive and base­line trend-fol­low­ing strate­gies, togeth­er with the draw­down of the adap­tive strat­e­gy. Both strate­gies com­bine fast-EWMA spans of 2, 4, 6, 10, 20, 35, 60, and 100 trad­ing days, while the adap­tive strat­e­gy in addi­tion adds the fast-EWMA span of 3. Returns are not com­pound­ed; start­ing cap­i­tal is 1 and the annu­al­ized risk tar­get is 25%.

Fig­ure 14 com­pares the equi­ty curves of adap­tive trend and the base­line mod­el. There are almost no extend­ed peri­ods in which the adap­tive strat­e­gy under­per­forms the base­line. It is slight­ly bet­ter most of the time.

Unfor­tu­nate­ly the adap­tive ver­sion also does not elim­i­nate the extend­ed draw­down that began in 2023, although it slight­ly reduces its mag­ni­tude.

Table 3 shows some per­for­mance sta­tis­tics over the whole back­test peri­od from July 1975 through July 2026. Table 4 com­pares the base­line strat­e­gy to the adap­tive ver­sion for the time­frame begin­ning Jan­u­ary 2008.

The com­par­i­son since 2008 is par­tic­u­lar­ly infor­ma­tive. The adap­tive ver­sion is still slight­ly bet­ter in near­ly every regard. But the improve­ment is small.

The result is encour­ag­ing rather than rev­o­lu­tion­ary. Across most look­backs, the adap­tive ver­sion adds approx­i­mate­ly 0.05 Sharpe-ratio points after costs. The two para­me­ter pro­files are almost par­al­lel.

Sta­tis­tic Val­ue
Return after Costs [% p.a.] 35.30
Sharpe 1.44
Max­i­mum Draw­down [%] 64.99
Avg Draw­down Dura­tion [days] 24.48
Max Draw­down Dura­tion [days] 1066
Month­ly Skew 0.41
Low­er Tail 1.78
Upper Tail 1.60
Win­ning Days [%] 55.14
Per­for­mance sta­tis­tics for the adap­tive trend-fol­low­ing strat­e­gy using the select­ed mix of look­backs, July 1975–July 2026
Sta­tis­tic Base­line Adap­tive
Return after Costs [%] 14.72 15.97
Sharpe 0.65 0.71
Max. Draw­down [%] 66.03 64.99
Avg DDD [days] 51.65 46.87
Max DDD [days] 1066 1066
Month­ly Skew 0.50 0.60
Low­er Tail 2.07 1.98
Upper Tail 1.61 1.62
Win­ning Days [%] 54.32 54.05
Com­par­i­son of per­for­mance sta­tis­tics of base­line trend ver­sus adap­tive trend using the select­ed mix of look­backs, Jan­u­ary 2008–July 2026

This is not the sort of improve­ment that pro­duces an excit­ing mar­ket­ing pre­sen­ta­tion. It is exact­ly the sort that can mat­ter in a sys­tem­at­ic research process: small, broad, eco­nom­i­cal­ly moti­vat­ed, and not depen­dent on one for­tu­nate para­me­ter.

Statistical Validation

The next phase of the research process is pri­mar­i­ly a val­i­da­tion step. The strat­e­gy is then sub­ject­ed to sev­er­al sta­tis­ti­cal pro­ce­dures such as sen­si­tiv­i­ty analy­sis of costs, delayed exe­cu­tion, walk-for­ward analy­sis, time sta­bil­i­ty tests, sig­nif­i­cance tests, and return sim­u­la­tions.

All sta­tis­ti­cal tests have to be inter­pret­ed with a grain of salt because finan­cial mar­kets are inher­ent­ly non-sta­tion­ary. Unfor­tu­nate­ly, no sta­tis­ti­cal test can prove that a strat­e­gy will con­tin­ue to work. They can, how­ev­er, inval­i­date a strat­e­gy.

The analy­sis of costs plays an impor­tant role in mak­ing sure, that a strat­e­gy does not only work in typ­i­cal con­di­tions but also in times of mar­ket stress. Dur­ing peri­ods of mar­ket stress not only liq­uid­i­ty may dry up but trad­ing may also become very expen­sive due to widen­ing of bid/ask spreads.

In a worst-case sce­nario you have to pay the full bid/ask spread each time you change a posi­tion in your port­fo­lio. That may occur due to enter­ing or clos­ing new posi­tions or if one future con­tracts get close to expiry and has to be rolled over. In addi­tion you have to pay com­mis­sions.

Ana­lyz­ing the costs of a stand-alone strat­e­gy is a con­ser­v­a­tive mea­sure, because costs typ­i­cal­ly come down due to less trad­ing if you com­bine sev­er­al strate­gies where some trad­ing sig­nals can­cel each oth­er out.

For the adap­tive trend strat­e­gy, the cost analy­sis is straight­for­ward: As already dis­cussed on Fig­ure 7, the costs for the strat­e­gy are about 0.15 Sharpe. A dou­bling of the costs to 0.3 Sharpe is also tol­er­a­ble over the full back­test but becomes crit­i­cal after 2008 where the per­for­mance drops.

The next check will be relat­ed to imple­men­ta­tion risk. If you run a ful­ly auto­mat­ed sys­tem­at­ic futures strat­e­gy there is always the risk of some­thing going wrong. The con­nec­tion to the bro­ker may get lost, pow­er out­ages, prob­lems at the bro­ker or sim­ply bugs in the soft­ware.

On a strat­e­gy lev­el, this results in orders that the sys­tems wants to do but where the sys­tem does not get a fill. With the adap­tive trend strat­e­gy orders will be gen­er­at­ed every day. If some orders of the pri­or day did not get filled the sys­tem will gen­er­ate it sig­nals as usu­al the next day, com­pare the want­ed posi­tions with the cur­rent port­fo­lio and place the need­ed orders.

If some­thing went wrong the pri­or day the miss­ing orders typ­i­cal­ly will be exe­cut­ed the next day (if the sig­nal stayed the same). Any errors will result in delayed exe­cu­tion. The strat­e­gy can be sub­ject­ed to dif­fer­ent delays in exe­cu­tion to get a feel­ing about its sen­si­tiv­i­ty to exe­cu­tion time of the orders.

Fig­ure 15 shows the loss of Sharpe ratio when delay­ing the orders of the strat­e­gy by up to 5 trad­ing days. As the strat­e­gy trades quite slow­ly the per­for­mance decrease is rel­a­tive­ly small with low sen­si­tiv­i­ty. That for sure does not pose a prob­lem if unre­li­able infra­struc­ture, soft­ware or the bro­ker does not exe­cute trades a few times a year. Nev­er­the­less, the graph also shows that you should not go on vaca­tion for an extend­ed peri­od of time with­out mon­i­tor­ing your sys­tem.

Decline of Sharpe ratio (after costs) of sys­tem­at­i­cal­ly delay­ing all orders of the strat­e­gy by 1 to 5 trad­ing days.

The next check will be con­cern­ing the sta­bil­i­ty of the mod­el para­me­ters over time with a walk-for­ward sim­u­la­tion. Think of this as forc­ing your younger self to make today’s deci­sions with­out know­ing tomorrow’s data. Once a year para­me­ters (like the mix of look­backs or the cal­cu­la­tion of turnover for each instru­ment to decide which look­backs to trade with each instru­ment) are recal­i­brat­ed and then trad­ed for a year.

This process is sim­u­lat­ed so that the recal­i­bra­tion process gets all data avail­able up to a that year, then a year­ly recal­i­bra­tion with extend­ed his­tor­i­cal data is done and the next year is sub­ject­ed to a back­test. The recal­i­bra­tion process starts in 1985 to at least have enough data for a few instru­ments and then eval­u­ates year by year.

The results can then be com­pared to the full sam­ple back­test of Table 1. The results are shown in Fig­ure 16.

Loss of Sharpe ratio (after costs) of year­ly recal­i­bra­tion using a walk-for­ward analy­sis vs the cal­i­bra­tion of the full sam­ple.

Inter­est­ing­ly the year­ly re-cal­i­bra­tion selects very sim­i­lar para­me­ters even in the begin­ning of the back­test. In lat­er years they are most­ly iden­ti­cal to the full-sam­ple back­test because the extend­ing win­dow of data sta­bi­lizes the results more and more.

On aver­age the walk-for­ward results are 0.05 Sharpe ratio points worse than the full in-sam­ple fit. The per­for­mance loss is so small, that this strat­e­gy has only minor esti­ma­tion risk.

Fig­ure 17 shows the sta­bil­i­ty of the Sharpe ratio of the strat­e­gy over time. We already dis­cussed ear­li­er the notice­able drop in per­for­mance begin­ning 2008 and the addi­tion­al drop in 2022. Should you be con­cerned about this drop in per­for­mance?

Rolling 1‑year Sharpe ratio of the adapt­ed trend strat­e­gy. Risk-free rate is 0%, no geo­met­ric com­pound­ing.

To answer this ques­tion, we boot­strap the Sharpe ratio over the full sam­ple, since 2008 and since 2022. We use a one-sided sta­tion­ary block-boot­strap to test, whether the annu­alised after-cost Sharpe ratio exceeds zero.

A block boot­strap does not draw indi­vid­ual dai­ly returns inde­pen­dent­ly. Instead, it resam­ples con­sec­u­tive blocks of returns. A mean block length of 20 trad­ing days is used. The block boot­strap has the advan­tage that the struc­ture of auto­cor­re­lat­ed returns (like in trend fol­low­ing) is pre­served. The block length of 20 was cho­sen to approx­i­mate­ly reflect the depen­dence hori­zon induced by the strategy’s typ­i­cal hold­ing peri­od.

This sta­tis­ti­cal tech­nique can also answer whether the expect­ed after-cost Sharpe ratio is pos­i­tive. (null hypoth­e­sis H0 Sharpe ratio zero or neg­a­tive, H1 pos­i­tive Sharpe) at the cho­sen sig­nif­i­cance lev­el of \(\mathsf{\alpha = 5}\)%. The result is shown in Table 5.

The esti­mat­ed Sharpe ratio dif­fers sub­stan­tial­ly across the three sam­ple peri­ods, sug­gest­ing that the strategy’s per­for­mance has changed over time. While the true Sharpe ratio of the full sam­ple lies between 1.13 and 1.70, the upper bound of the Sharpe ratio of the sam­ple begin­ning 2022 comes in at 0.99.

The con­fi­dence inter­vals indi­cate a sub­stan­tial dete­ri­o­ra­tion in esti­mat­ed risk-adjust­ed per­for­mance after 2022. While the full sam­ple sup­ports a strong­ly pos­i­tive Sharpe ratio, the post-2022 sam­ple no longer pro­vides suf­fi­cient evi­dence that the expect­ed Sharpe ratio exceeds zero.

Peri­od Sharpe Low­er Upper H0
1975–2026 1.44 1.13 1.70 reject
2008–2026 0.71 0.25 1.21 reject
2022–2026 -0.02 -1.13 0.99 accept
Boot­strapped Sharpe ratios of the dai­ly returns of the adap­tive trend strat­e­gy. The boot­strap used is a sta­tion­ary block boot­strap with a mean block length of 20. 10,000 sim­u­la­tions per peri­od with a con­fi­dence lev­el of 95%. Sharpe ratios are non com­pound­ing, annu­al­ized with a risk-free rate of 0%.

The sam­ple begin­ning in 2022 even accepts the null hypoth­e­sis (that the Sharpe ratio is zero or neg­a­tive) at the sig­nif­i­cance lev­el of 5%. It there­fore can­not infer that the per­for­mance is pos­i­tive.

All this does not inval­i­date the strat­e­gy but should raise some con­cerns. A strat­e­gy like this might get a place in the port­fo­lio but must not be over­weight­ed for sure.

The last sta­tis­ti­cal test is con­cerned with the cre­ation of the strat­e­gy. As men­tioned ear­li­er, it is not advis­able to test many dif­fer­ent sig­nal designs. You may just arrive at a good sig­nal by chance. Of course this can­not be com­plete­ly avoid­ed, because you have to test some­thing.

Dur­ing strat­e­gy devel­op­ment, sev­er­al eco­nom­i­cal­ly moti­vat­ed sig­nal spec­i­fi­ca­tions were eval­u­at­ed. Select­ing the best-per­form­ing spec­i­fi­ca­tion after observ­ing the his­tor­i­cal data intro­duces data-snoop­ing bias because the final mod­el is cho­sen from a larg­er can­di­date set.

How did that process work in the cre­ation of the adap­tive trend strat­e­gy?

In total I test­ed four dif­fer­ent sig­nal designs. The main idea was to add the adap­tive look­back to the strat­e­gy. My first instinct was to short­en both mov­ing aver­ages. It bare­ly changed the results. That fail­ure turned out to be use­ful because it sug­gest­ed that the prob­lem was not react­ing too slow­ly to new infor­ma­tion, but for­get­ting old infor­ma­tion too slow­ly. To check the results, I also added it exclu­siv­ly to EWMA-fast (also had no mate­r­i­al effect the result).

Con­ven­tion­al sig­nif­i­cance tests eval­u­ate only the final select­ed mod­el and there­fore ignore the fact that sev­er­al com­pet­ing spec­i­fi­ca­tions were explored. Hansen’s SPA test explic­it­ly accounts for this research-selec­tion process by eval­u­at­ing all can­di­date strate­gies joint­ly.

The null hypoth­e­sis states that none of the test­ed spec­i­fi­ca­tions pos­sess­es pos­i­tive pre­dic­tive abil­i­ty after account­ing for mod­el selec­tion. Unlike ordi­nary sig­nif­i­cance tests applied only to the final strat­e­gy, the SPA test adjusts for the fact that sev­er­al com­pet­ing spec­i­fi­ca­tions were con­sid­ered dur­ing the research process. The results are shown in Table 6.

Sig­nal Sharpe
Base­line 1.37
Adap­tive Fast+Slow 1.37
Adap­tive Fast 1.35
Adap­tive Slow 1.44
SPA p‑value 0.028
Sharpe ratios of dif­fer­ent research can­di­dates of the adap­tive trend strate­gies with p‑value of Hansen’s SPA test.

Portfolio Integration

The last step in the research work­flow is to incor­po­rate the devel­oped strat­e­gy into a port­fo­lio of strate­gies. A sim­ple port­fo­lio can be con­struct­ed using fixed port­fo­lio weights or fixed risk allo­ca­tions.

Anoth­er pos­si­bil­i­ty is a more quan­ti­ta­tive approach using some form of port­fo­lio opti­miza­tion. Because this is a com­plex top­ic in its own right, it is not cov­ered in this arti­cle and we stick to sim­ple port­fo­lio weights for now.

Each strat­e­gy is cat­e­go­rized as either “trend” or “mean rever­sion”. Strate­gies of the trend­ing type prof­it from diverg­ing prices (either grow­ing or falling) while mean revert­ing strate­gies prof­it from prices return­ing to some mean val­ue. These two types have very dif­fer­ent cor­re­la­tions among each oth­er.

The cat­e­go­riza­tion is done to have some fin­er con­trol over the num­ber of strate­gies in each cat­e­go­ry. The trend cat­e­go­ry has a weight of 60% and the mean rever­sion cat­e­go­ry of 40%. With­in each cat­e­go­ry the strate­gies are equal­ly weight­ed. If “trend” con­sists of 5 strate­gies, then each strat­e­gy gets a final weight of 12% (60% x 20%).

The cor­re­la­tions between all strate­gies of the port­fo­lio are shown in Fig­ure 18.

Cor­re­la­tions of returns between dif­fer­ent strate­gies. Each strat­e­gy can be a fam­i­ly of para­me­ters.

There are two dis­tinc­tive clus­ters of cor­re­la­tions between the strate­gies of type trend and of type mean rever­sion observ­able. Each clus­ter has high­er cor­re­la­tions with­in the clus­ter. The adap­tive trend strat­e­gy also has high cor­re­la­tions with three oth­er trend strate­gies of over 90%.

Nev­er­the­less each strat­e­gy plays its role in the port­fo­lio like an ant in an ant colony. This par­tic­u­lar ant (adap­tive trend) increas­es the Sharpe ratio of the whole port­fo­lio by a small amount of 0.08. Not great, but also nice to have!

Conclusion

We set out to build a trend-fol­low­ing strat­e­gy from first prin­ci­ples, and first prin­ci­ples deliv­ered — just not a mir­a­cle. The eco­nom­ic ratio­nale held up under scruti­ny. The sig­nal behaved itself. The para­me­ter sur­face was smooth, and the walk-for­ward tests were bor­ing, which in this line of work is the high­est com­pli­ment a num­ber can receive. All this back and forth pro­duced a mod­i­fied strat­e­gy bet­ter by about 0.05 Sharpe-ratio points by itself and 0.08 to the port­fo­lio as a whole.

That is, by any hon­est mea­sure, not very much. It won’t head­line a pitch deck. It won’t make anyone’s year. What it will do is sit qui­et­ly in a port­fo­lio next to a dozen oth­er unglam­orous improve­ments — each one small enough to look like round­ing error, each one earned the hard way: a ratio­nale writ­ten down before a sin­gle back­test ran, a sig­nal checked for spikes before it was checked for prof­it, and a fix that sur­vived a test built specif­i­cal­ly to catch researchers fool­ing them­selves.

That, in the end, is what craft­ing a strat­e­gy actu­al­ly looks like. Not a light­ning bolt of insight, but an ant colony’s worth of small, unglam­orous con­tri­bu­tions — most too mod­est to notice on their own, all mov­ing in rough­ly the same direc­tion, and every one of them still run­ning in cir­cles some­where, wait­ing for the next iter­a­tion to prove it wrong.

Footnotes

  1. In this arti­cle a future trad­ing a par­tic­u­lar under­ly­ing asset is called an instru­ment. There are cas­es where the same asset is trad­ed as two dif­fer­ent instru­ments, like trad­ing two dif­fer­ent sizes of the same under­ly­ing (like mini and micro con­tracts).↩︎

  • Crafting a Trading Strategy

    Developing a systematic trading strategy rarely begins with a perfect idea. Instead, it is an iterative process of forming hypotheses, testing them, discarding what does not work, and refining what does. In this article, we develop a futures trading strategy… Read more.

  • In Search of Diversification by Selecting Strategy Parameters

    Diversification is one of the most important tools for a systematic trader. This article will dive into the diversification possibilities of one simple strategy via different parameters. The article will both discuss the theoretical foundations and apply these to a… Read more.