﻿<?xml version="1.0" encoding="utf-8"?><Search><pages Count="168"><page Index="1" isMAC="true"><![CDATA[CONTRIBUTIONS IN STATISTICS AND INFERENCECelebrating Nazaré Mendes Lopes' Birthday]]></page><page Index="2" isMAC="true"><![CDATA[Comissa˜o Editorial/Editorial BoardAna Paula SantanaJu´lio Severino NevesMaria Paula Martins Serra de OliveiraFinancial support for the edition of this volume of Textos de Matem´atica is gratefully acknowledged to Centre for Mathematics of the University of Coimbra (CMUC) — UID/MAT/00324/2013, funded by the Portuguese Government through FCT/MEC and co-funded by the European Regional Development Fund through the Partnership Agreement PT2020.]]></page><page Index="3" isMAC="true"><![CDATA[CONTRIBUTIONS IN STATISTICS AND INFERENCECelebrating Nazare´ Mendes Lopes’ BirthdayEsmeralda Gon¸calves Paulo Eduardo Oliveira Carlos Tenreiro (editors)Textos de Matem´atica/Mathematical TextsVolume 47Departamento de Matem´atica da Universidade de Coimbra Portugal 2015]]></page><page Index="4" isMAC="true"><![CDATA[Title: Contributions in Statistics and Inference: Celebrating Nazar´e Mendes Lopes’ BirthdayEditors: Esmeralda Gonc¸alves, Paulo Eduardo Oliveira, and Carlos TenreiroPublisher: Departamento de Matem´atica, Universidade de Coimbra/Mathematics Department of the University of Coimbra, PortugalCollection: Textos de Matem´atica/Mathematical Texts Number: 47DL: 391956/15ISBN: 978-972-8564-51-3Camera-ready copy prepared using LATEX Printed and bounded by Tipografia Macasi, Lda Printed in Coimbra, Portugal]]></page><page Index="5" isMAC="true"><![CDATA[]]></page><page Index="6" isMAC="true"><![CDATA[Nazar´e Mendes Lopes]]></page><page Index="7" isMAC="true"><![CDATA[FOREWORDThis special issue of Textos de Matem´atica is dedicated to our colleague and friend Nazar´e Mendes Lopes. It includes extended versions of some of the talks presented at the Workshop on Statistics and Inference held at the Department of Mathematics of the University of Coimbra on the occasion of her sixtieth birthday, and some other contributions by researchers invited to participate in this volume.Nazar´e Mendes Lopes is one of the founding members of the Probability and Statistics group at the University of Coimbra. Following an e↵ort by the Department of Mathematics to modernize the teaching of probability and sta- tistics and to foster research activities in this field, Nazar´e Mendes Lopes found herself in close connection with the French school that was guiding these e↵orts. After her graduation in Applied Mathematics by the University of Coimbra, she undertook post-graduate studies in probability and statistics supervised by Raymond Moch´e, later leaving to Paris to work with Jean Ge↵roy to finally conclude her PhD in the very beginning of 1985 at the University of Coimbra.Nazar´e Mendes Lopes early research interests concentrated on point pro- cesses and nonparametric methods. Shortly after this initial period, Nazar´e Mendes Lopes redirected her interests to the domain of time series. This has re- mained her main research area for which Nazar´e Mendes Lopes has contributed to the establishment of fundamental properties and the comprehension of the structure of some time series models, particularly bilinear and conditionally heteroscedastic ones.Nazar´e Mendes Lopes teaching qualities have always been well recognised. The enthusiasm and commitment she puts in her teaching has certainly helped her e↵orts in attracting young students to probability and statistics. This enthu- siasm and seriousness are reflected in the published monographs, that evolved from classroom notes, where a particular, rather mathematical, view of the field is explained.Finally, we would like to thank the special presence of Jean Michel Zakoian, Filipa Silva, Kamil Feridun Turkman and Maria Ivette Gomes in the workshop, the authors who contributed to this volume, the reviewers for their valuable col- laboration, the editorial board of Textos de Matema´tica, Maria Paula Oliveira, Ana Paula Santana and Ju´lio Neves, for supporting and promoting the publi- cation of this special issue, and the Centre for Mathematics of the University of Coimbra for its financial support. Without them the workshop and the pub- lication of this volume would not have been possible.Coimbra, February 2015The editors]]></page><page Index="8" isMAC="true"><![CDATA[]]></page><page Index="9" isMAC="true"><![CDATA[M. I. GomesCONTENTSNAZARE´ and ARCH processes: extremal index estimation 1H. FerreiraMax-min dependence coe cients for Multivariate Extreme Value Distributions 13E. Gon¸calves and C. M. MartinsThe Taylor property in non-negative autoregressive and bilinear stochastic processes 25J. LeiteOn the ability of the power threshold GARCH model to capture the Taylor e↵ect 37M. M. NevesBootstrap and Jackknife methods in extremal index estimation: a review 49P. E. OliveiraMean square error in regression estimation with functional data 67 I. Pereira, M. Scotto, and R. NicoletteInteger-valued self-exciting periodic threshold autoregressive processes 83A. C. Rosa and M. E. NogueiraA note on kernel estimation of the conditional quantile function in continuous time ergodic processes 95M. E. SilvaModelling time series of counts: an INAR approach 107M. G. TemidoOn the extremes of stationary gaussian random fields under strong dependence 123P. de Zea Bermudez, M. A. Amaral Turkman, and K. F. TurkmanParameter Estimation of Bilinear Processes using Approximate Bayesian Computation 135]]></page><page Index="10" isMAC="true"><![CDATA[]]></page><page Index="11" isMAC="true"><![CDATA[NAZARE´ AND ARCH PROCESSES: EXTREMAL INDEX ESTIMATIONM. IVETTE GOMESTo my good friend Nazar´e Mendes-Lopes, a token of friendshipAbstract. AfterafewmemoriesrelatedtomyfriendNazar´e,definitions of extremal index and ARCH processes are put forward. A brief revision of classical and corrected bias generalized jackknife estimation of the extremal index is further considered. Finally, we dedicate a short and personal tribute to Nazar´e, speaking about a possible co-operation, ‘the heart of Science’.1. Memories and scientific scope of the articleAs far as I remember, I met Nazar´e for the first time thirty years ago, in Lagos, Algarve, 1984, during the “III Col´oquio de Estat´ıstica e Investiga¸c˜ao Operacional”. Nazar´e spoke on recent results she had already published at Publications de l’Institut de Statistique de l’Universit´e de Paris (Mendes-Lopes, 1983, 1984), and it is still quite vivid in my memory her talk on what I used to call ‘marked point processes’. I have indeed found quite interesting the French nomenclature, ‘chromatiques’, for this type of processes, a terminology that I have never forgotten. She had been working for M.Sc. in France, at the Uni- versit´e de Paris VI (Pierre et Marie Curie), having begun there her Ph.D. project under the supervision of Jean Ge↵roy, a pioneer in the field of extreme value theory (EVT), my elected field of research. Despite of having got a de- gree of ‘Docteur de Troisi`eme Cycle’, in Paris, 1982, she had to prepare a Ph.D. thesis in Portugal. And I had the honor to be in ‘Sala dos Capelos’, Universi- dade de Coimbra, for the first time in 1985, to discuss her Ph.D. thesis. I also vividly recall the table where we had to place our documents, small enough for Professor Jean Ge↵roy only . . .Accepted: 17 February 2015.2010 Mathematics Subject Classification. Primary 60G70; Secondary 60G10, 60G17. Key words and phrases. ARCH processes, extremal index, friendship, generalized jackknife,statistics of extremes.The work was supported by National Funds through FCT (Funda¸c˜ao para a Ciˆencia e aTecnologia), PEst-OE/MAT/UI0006/2014 project.1]]></page><page Index="12" isMAC="true"><![CDATA[2 M. I. GOMESA few years later my contacts with Coimbra, and also with Nazar´e, whom I already considered a very good friend at the time, were deepened, namely be- cause Nazar´e asked me to supervise Helena Ferreira (my good friend Lena), who was an Assistant at the Department of Mathematics, University of Coimbra, at the time. In Lisboa, DEIOC (Departamento de Estat´ıstica, Investiga¸c˜ao Operacional e Computa¸c˜ao) already had an M.Sc. Degree in Statistics and Operations Research, but in Coimbra, Lena had to be submitted to what we used to call PAPCS (‘Provas de Aptid˜ao Pedag´ogica e Capacidade Cient´ıfica’). Her monograph, entitled ‘Valores Extremos em Esquemas de Dependˆencia Fraca’, was discussed in 1989, and with this supervision my knowledge of ex- tremes of dependent schemes has improved a lot, despite of the fact that I have never worked deeply under such frameworks. Surely, we learn a lot with most of our students, something I miss a bit now that I am retired. And I had bright students working under my supervision not only in the aforementioned topic but also in other topics, among whom I mention Lena, who got a Ph.D. in 1994 (Ferreira, 1994).Later on, Gra¸ca (Maria da Gra¸ca Temido), another good friend and now Assistant Professor at Universidade de Coimbra, came to Lisboa for M.Sc., at the time on Probability and Statistics, and at DEIO (Departamento de Estat´ıstica e Investiga¸c˜ao Operacional). Gra¸ca worked on Extreme Values in a Normal Model and finished her M.Sc. degree in 1992, under my supervision. She went on for Ph.D., under the supervision of Lu´ısa Canto e Castro and myself, having got her degree in 2000 (Temido, 2000). But despite of having had no more Ph.D. students from Coimbra my contacts and friendship with several colleagues from Coimbra have always been very gratifying.When trying to find scientific research links with Nazar´e’s work, I could essentially find two topics in which we both have independently done research, Statistical Quality Control and ARCH/GARCH processes. Since I know that Nazar´e has worked with dependent processes for more than 20 years, as can be seen from two recent papers (Gon¸calves and Mendes-Lopes, 2013; Gonc¸alves et al., 2013), and two old ones (Brito et al., 1992; Gon¸calves and Mendes-Lopes, 1993), among more than two handfuls of published articles on this type of dependent processes, I have thus decided to choose the second topic, where I have only developed some work essentially related to the extremal index in ARCH processes. In Section 2 of this article, the notion of extremal index is provided. Section 3 is dedicated to the definition and a few properties of the ARCH processes. In Section 4, a brief review of some EI-estimators is performed. Finally, Section 5 is dedicated to a short, personal tribute and a possible challenge to Nazar´e.]]></page><page Index="13" isMAC="true"><![CDATA[NAZARE´ AND ARCH PROCESSES: EXTREMAL INDEX ESTIMATION 32. The extremal indexLet {Xn}n 1 be a stationary sequence from an underlying model F, under adequate asymptotic dependence conditions. Let {Yn}n 1 be the associated independent, identically distributed (IID) sequence, from the same underlying model F. Let us further denoteMnX ⌘ Xn:n := max(X1,...,Xn) and MnY ⌘ Yn:n := max(Y1,...,Yn). If there exist sequences of constants {an > 0} and {bn 2 R} such thatlim P ✓ MnY   bn  x◆ = G(x), with G non-degenerate, n!1 anthen (Gnedenko, 1943)G(x)⌘EV⇠(x)=⇢ exp  (1+⇠x) 1/⇠ ,1+⇠x>0, if ⇠6=0,exp(  exp( x)), x 2 R, if ⇠ = 0,is the well-known extreme value (EV) cumulative distribution function (CDF), being ⇠ the extreme value index (EVI).Let us next think on the stationary sequence {Xn}n 1. If the conditions which enable us to guarantee the existence of an extremal index (EI), denoted by ✓, held (Leadbetter et al., 1983), thenlim P✓MnX  bn x◆=EV✓⇠(x), 0✓1. n!1 anIndeed, we shall exclude the ‘slight pathological’ case ✓ = 0.Remark 2.1. Note that the EV⇠ CDF is max-stable, and consequently,EV✓⇠(x)=EV⇠✓x  ✓◆,  ✓=✓⇠ 1,  ✓=✓⇠.  ✓ ⇠To better understand the intuitive meaning of the EI, let us think on the point process of exceedances over high thresholds. For any ⌧ > 0, let un = un(⌧), n   1, be a level such thatF(un)=1 ⌧/n+o(1/n), asn!1, (2.1)a so-called normalized level, i.e. a level such that n(1 F(un))!⌧>0, or equivalently, a level such that F n (un ) ! exp( ⌧ ), as n ! 1. Note that it is possible to choose such a level for any CDF such that (1   F (x ))/(1   F (x)) ! 1, as x ! 1, which will be assumed throughout. Then, with IA denoting the indicator function of A,Y XnSn = I{Yj>un} j=1]]></page><page Index="14" isMAC="true"><![CDATA[4 M. I. GOMESis Binomial(n, 1   F (un )), and consequently converges towards a Poisson(⌧ ),as n ! 1, whereasXnSnX = I{Xj>un}j=1converges, as n ! 1, towards a compound Poisson. Then P(MnY  un) = P(SnY = 0) ! e ⌧, whereas P(MnX  un) = P(SnX = 0) ! e ✓⌧, as n!1. This means that the intensity ⌧ of the Poisson limiting exceedances, in the IID case, becomes ✓⌧, i.e. the point process limit for the time normalized upcross- ings of high levels is also a Poisson point process but with intensity ✓⌧. Under independence (or even adequate quite weak dependence), this point process converges to a homogeneous Poisson process, as n ! 1, but when there is a slightly stronger local dependence, clusters of exceedances may occur and the limiting process of exceedances may be a compound Poisson process. Indeed, for a large class of weakly dependent processes, an upcrossing is generally followed by a cluster of exceedances and therefore the clusters may be roughly identified by the occurrence of upcrossings. Indeed, Leadbetter and Nandagopalan (1989) proved that the EI can then also be defined as the reciprocal of the ‘mean time of duration of extreme events’, and it is directly related to the exceedances of high levels. We have✓ = 1 = lim P(X2 un|X1 >un) limiting mean size of clusters n!1= lim P(X1  un|X2 > un), n!1where un is a sequence of values such that (2.1) holds. Then,P(Yn:n  un) = Fn(un)  ! e ⌧ and P(Xn:n  un)  ! e ✓⌧.n!1 n!1Remark 2.2. The limiting distribution of normalized maximum values of both {Xn} and {Yn} is thus of the same type, but there exists a ‘shrinkage’ of maximum values. This leads to ‘clusters of exceedances of high levels’ with a mean size greater than 1. To illustrate these facts, we next provide in Figure 1, sample paths associated with{2.1} an IID Fr´echet sequence, {Yi}i 1, from F(x) = exp{ x ↵}, x   0, ↵ = 2 (⇠ = 1/2 = 0.5, ✓ = 1),{2.2} a stationary 2-dependent sequence also with Fr´echet margins and ↵ = 2, defined asXi = 2 1/↵ max (Yi, Yi+1) , Yi given in {2.1}, i   1, for which ✓ = 0.5, and]]></page><page Index="15" isMAC="true"><![CDATA[NAZARE´ AND ARCH PROCESSES: EXTREMAL INDEX ESTIMATION 5{2.3} an autoregressive for maxima (ARMAX) Fr´echet stationary sequence, Xi+1 = max(Xi,Ui), Ui :=(  ↵  1)1/↵ Yi,with ↵ = 2, {Yi}i 1 given in {2.1} as before, and with   = 0.8. Then, the EI is ✓ = 1    1/⇠ = 0.36 (see Alpuim, 1989, and Hall, 1996).7766554433221100 76 5 4 3 2 1 0Figure 1. Sample paths of the processes in {2.1} (top left, ✓ = 1), {2.2} (top right, ✓ = 0.5) and {2.3} (bottom, ✓ = 0.36), all from the same underlying F.3. The ARCH processThe autoregressive conditional heteroscedastic (ARCH) process introducedby Engle (1982) is defined byVn=Wnq + Vn2 1, n 1, (3.1)where {Wn} are IID standard normal random variables (RVs),   > 0, 0< <1.Then, Xn =Vn2 issuchthatXn =AnXn 1 +Bn, n 1, X0  0, (3.2)An =  Wn2, Bn =  Wn2, {(An, Bn), n   1} IID R2+-valued random pairs. The stochastic di↵erence equation in (3.2) was introduced by Kesten (1973) and further studied in Vervaat (1979), where several examples are provided, like for instance the stock of material checked at regular time intervals, with An the intrinsic decay or increase of the stock and Bn the quantity added or taken]]></page><page Index="16" isMAC="true"><![CDATA[6 M. I. GOMESaway just before time n. The extremal behaviour of such a stochastic di↵erence sequence was studied in de Haan et al. (1989).The most striking feature of the ARCH process in (3.1) lies in the fact that although the building RVs, {Wn}, are standard normal, in the max-domain of attraction of the Gumbel law, ⇤(x) := exp( exp( x)), x 2 R, with a light exponential tail (⇠ = 0), Vn has heavy Pareto-like right-tails, i.e. ⇠ > 0.Let us consider the process in (3.2). If there is a positive real ↵ such that EA↵n=1, EA↵nln+An<1, 0<EBn↵<1,the process Xn, in (3.2), has an extremal index ✓ given by Z1 Yj !✓ = ↵ P sup Ai  y 1 y ↵ 1dy. 1 j 1 i=1Denoting by Z a strict Pareto(↵) RV, with CDF FZ (z) = 1   z ↵, z   1, and noticing that ✓ = E(IA), with ( Yj )A=Zsup Ai1,j 1 i=1de Haan et al. (1989) have simulated the value of the EI of |Vn| for di↵erent val- ues of  . In their simulation they have not used the event A, but the equivalenteventj 1 i=1 j 1 i=1 where E is a standard exponential RV.Regarding the EI of the ARCH process, one can use the fact thatVn =d Cn pVn2 , (3.3)with {Cn} IID RVs, independent of Vn, and such that, for all n   1, P(Cn =  1) = P(Cn = +1) = 1/2. In Table 1 we show a few simulated values of ✓ for Vn and |Vn|.Table 1. EI simulated values for the absolute values of the ARCH process and the ARCH process.(Xj )(EXj ) B= lnY+sup lnAi0 = ↵+sup lnAi0 , 0.10.30.50.70.9✓(Vn)0.9990.9390.8350.7210.612✓(|Vn|)0.9970.8870.7270.5790.460]]></page><page Index="17" isMAC="true"><![CDATA[NAZARE´ AND ARCH PROCESSES: EXTREMAL INDEX ESTIMATION 7 In Figure 2, we picture a sample path of Vn2, Vn given in (3.1), with   = 0.68and   = 0.04. Like this we get for Vn, ↵ = 2 and ✓ = 0.84.7 6 5 4 3 2 1 0Figure 2. Sample path of the square of an ARCH sequence, V 2 =Wn2  Vn2 1 +  , =0.68,  =0.04 (↵=2, ✓=0.84).nWe further mention that the finite joint structure of the extremes of an ARCH process enabled Gomes et al. (2004) to identify a peculiar ‘unexpected’ phenomenon, in the sense that the EI, one of the most relevant parameters of extreme events, seems not to be the adequate object to consider, when we need to infer on failure probabilities during a finite future time interval (for further details see Gomes et al., 2004, and Gomes et al., 2006, for a short correction).4. Extremal index estimation4.1. Classical EI-estimators. Given a sample (X1, . . . , Xn) and chosen a suitable threshold u, a possible estimator of ✓ (Leadbetter and Nandagopalan, 1989) is given byn 1 I n 1 I PPj=1 [Xj >u,Xj+1u] j=1 [Xj u<Xj+1]ˆN ˆN✓n=✓n(u):= Pn = Pn.j=1I[Xj >u]j=1I[Xj >u]To have consistency, the high level u must be such that n(1 F(un)) = cn⌧ = ⌧n, ⌧n ! 1 and ⌧n/n ! 0 (Nandagopalan, 1990). But whenever dealing with EVI-estimation we often consider an intermediate sequence kn, i.e. a sequence k=kn suchthatkn !1,butk/n!0,asn!1.Suchakn-sequencehas been replaced, in an EI-estimation, by the sequence ⌧n = cn⌧ with cn ! 1 as n ! 1. To make the semi-parametric EI-estimation closer to the most common semi-parametric EVI-estimation, and denoting by {Xi:n}1in the sample of]]></page><page Index="18" isMAC="true"><![CDATA[8 M. I. GOMESascending order statistics (OSs) associated with the original sample {Xi}1in, it is thus sensible to consider a deterministic level u 2 [Xn k:n, Xn k+1:n) andthe estimatorˆN 1 nX 1✓n (k) := k I[Xj Xn k:n<Xj+1]. (4.1)j=1The EI-estimator is then a function of k, the number of OS’s higher than thechosen threshold. We further assume a sensible structure for the asymptotic bias, given byBiasn✓ˆN(k)o=' (✓)✓k◆+' (✓)✓1◆+o✓1◆+o✓k◆, (4.2) n1n2kknas n ! 1, and for any intermediate k (see Gomes et al., 2008). Indeed, for IIDdata (✓ = 1):Moreover, for ARMAX processes, in {2.3}, we getE n✓ˆN(k)o = ✓   ✓ ✓(✓ + 1) ✓ k ◆   3   2 ✓ ◆ (1 + o(1)).In the semi-parametric EI-estimation we have thus to cope with problems similar to the ones appearing in the EVI-estimation: increasing bias, as the threshold decreases and a high variance for high thresholds. It is then sensible to ask whether it is possible to improve the performance of estimators through the use of resampling methods. We can surely use the generalized jackknife (GJ) methodology, to reduce the bias of the EI-estimators, in (4.1). Indeed, in statistics we often put the question whether the combination of information can improve the quality of estimators of a certain parameter or functional. And the jackknife or GJ are resampling methodologies, that usually give a positive answer to such a question.4.2. Resampling methodologies and corrected-bias EI-estimation. The use of resampling methodologies has revealed to be promising in the esti- mation of the tuning parameter k, and in the reduction of bias of any estimator of a parameter of extreme events. Regarding bias, and due to the fact that at optimal levels, in the sense of minimal mean square error (MSE) we still have a non-null asymptotic bias, we are lead to use the GJ methodology. It is then enough to consider an adequate set of estimators of the parameter of extreme events under consideration, and to build a reduced-bias a ne combination of them. In Gomes et al. (2000, 2002, 2013), also among others, we can find an application of this technique to the EVI-estimation. To illustrate here the useEn✓ˆN(k)o=1+✓ 1  k◆(1+o(1)). n2knn2n2k]]></page><page Index="19" isMAC="true"><![CDATA[NAZARE´ AND ARCH PROCESSES: EXTREMAL INDEX ESTIMATION 9of these methodologies in EVT, we apply the GJ methodology to the afore- mentioned EI-estimator in (4.1), as performed in Gomes et al. (2008).Since the bias term of the aforementioned classical EI-estimator reveals two main components of di↵erent orders, as can be seen in (4.2), we need to use an a ne combination of three EI-estimators, i.e. an order-2 GJ-statistic. Let X = (X1,...,Xn) be a sample from F, and let Tn = Tn(X,F) be an estimator of a functional ✓(F), or of a parameter ✓. If the bias of our estimator reveals two main terms that we would like to remove, the GJ methodology advises us to deal with three estimators with the same type of bias.E nT (i)   ✓o = d (✓) '(i)(n) + d (✓) '(i)(n), i = 1, 2, 3, n1122the GJ-statistic (of order 2) is given by   T(1) T(2) T(3)       1 1 1    1Definition 4.1. Given three estimators of ✓, Tn , Tn and Tn , such that(3)    n    1 1 1         with ||A|| denoting, as usual, the determinant of the matrix A. Straightforwardly, one may state the following result.Proposition 4.2. Under the validity of (4.3), the statistic TGJ, defined in n(4.4), is unbiased for the estimation of ✓.The variance of the statistic TGJ is always larger than the variance of theoriginal estimators, but the MSE of TGJ is often smaller than that of any of(i)the statistics Tn , i = 1, 2, 3.   n n n       (1) (2)TGJ :=   '(1) '(2) '(3)       '1 '1 '1    , (4.4)   '(1) '(2) '(3)       '(1) '(2) '(3)    222 222nnGiven the information on the bias of the extremal index estimator ✓ˆN(k), nin (4.1), as stated in (4.2), and with bxc denoting the integer part of x, let us consider, just as in Gomes et al. (2008), the levels k, b kc + 1 and b 2kc + 1, dependent of a tuning parameter  , 0 <   < 1, and the class of estimators,( 2 + 1) ✓ˆN (b kc + 1)     ⇣✓ˆN  b 2kc + 1  + ✓ˆN(k)⌘ ✓ˆ G J (   ) ( k ) : = n n n .n (1  )2Among the members of this class, the aforementioned authors have been heuris-(1) (2) (3)(4.3)tically led to the choice   = 1/4, and to the EI-estimator,✓ˆGJ(k) := ✓ˆGJ(1/4)(k). (4.5) nn]]></page><page Index="20" isMAC="true"><![CDATA[10 M. I. GOMESA comparison of the N and the GJ EI-estimators, respectively given in (4.1) and (4.5), based on small-scale Monte Carlo techniques, led us to the following conclusions:(1) The ‘naive’ GJ estimator of ✓ in (4.5) exhibits for all simulated models and for all values of   stable sample paths as functions of k, as illus- trated in Figure 3, for samples of size n = 5000 from ARCH structures with ✓ = 0.1, ✓ = 0.5 and ✓ = 0.9.(2) For low values of  , to which correspond high values of ✓, theGJ EI-estimator has, at its optimal level, a smaller MSE than theN EI-estimator, at the expenses of the use of a larger number of top1.5values of n. 1.31.10.90.70.5 450000.1 4 5000-0.1 04 -0.3 5000OS’s. When   increases such an advantage no longer holds for small0.30.1 0.10.5 0.50.90.9! ! 0.9! ! 0.5! ! 0.110002000 3000 4000 5000k6000Figure 3. GJ EI-estimators, in (4.5), for samples of size n = 5000 -0.5from ARCH processes.5. A challenge and a tribute to Nazare´And next goes a challenge to Nazar´e: As I believe that “co-operation is the heart of Science”, we have never co-authored any article, but we have some similar research interests, I hope we can collaborate in the near future in some topic. And I am in particular thinking on the topic I suggested to Nazar´e at her Habilitation Degree: Linking the Choice of the Window in a Kernel Density Estimation with the Choice of the Threshold in Statistics of Extremes. Indeed, the tuning or nuisance parameter h = hn, the size of the window in a kernel density estimation, with n the sample size, needs to be such that hn ! 0 and nhn ! 1, as n ! 1. In statistics of univariate extremes, the crucial tuningˆJ!n (k)]]></page><page Index="21" isMAC="true"><![CDATA[NAZARE´ AND ARCH PROCESSES: EXTREMAL INDEX ESTIMATION 11parameter is the number k = kn of order statistics that should be used for a functional consistent estimation of parameters of extremes events, like the EVI. Andweneedtohavekn !1andkn/n!0,asn!1...And the challenge is followed by a short tribute: When I first met Nazar´e I immediately understood that I had met a special person. And I believed that it would be a forever friendship . . . . We have not had many chances to meet, but when we meet I always feel renovated and happy. Dinis, my husband for more than forty years, says that Nazar´e is not human, just because she cannot lie . . . I indeed think that this is the source of her rightness and fairness. On the other hand, we both think that Nazar´e is one of our more ‘human’ friends, in the sense that she deeply sympathises with other people feelings and problems, and that she feels sad and ashamed with poverty and injustice. Selma Lagerl¨of once wrote that true greatness is greatness of the heart, and I am happy to state that aside from her scientific achievements Nazar´e is great, a great human being and a great friend. And she enjoys a good laugh, a very enjoyable trait, since humans are much more interesting than saints.AcknowledgmentsThis research was partially supported by National Funds through FCT (Fundac¸˜ao para a Ciˆencia e a Tecnologia), through the project PEst-OE/MAT/UI0006/2014.References[1] M. T. Alpuim, An extremal markovian sequence, J. Appl. Probab. 26, 219–232, 1989.[2] M. Brito, E. Gonc¸alves, and N. Mendes-Lopes, Contribuic¸a˜o para o estudo de modelos heterosced´asticos, in: Estudos de Matem´atica em Homenagem ao Prof. Renato PereiraCoelho, Departamento de Matem´atica, Universidade de Coimbra, 131–144, 1992.[3] R. F. Engle, Autoregressive conditional heteroscedastic models with estimates of thevariance of United Kingdom inflation, Econometrica 50, 987–1007, 1982.[4] H. Ferreira, Condi¸c˜oes de Dependˆencia Local em Teoria de Valores Extremos, Tese de Doutoramento, Coimbra: Departamento de Matem´atica, Universidade de Coimbra, 1994.[5] B. V. Gnedenko, Sur la distribution limite du terme maximum d’une s´erie al´eatoire,Annals of Mathematics 44, 423–453, 1943.[6] M. I. Gomes, M. J. Martins, and M. Neves, Alternatives to a semi-parametric estimatorof parameters of rare events—the Jackknife methodology, Extremes 3 (3), 207–229,2000.[7] M. I. Gomes, M. J. Martins, and M. Neves, Generalized jackknife semi-parametric esti-mators of the tail index, Port. Math. (N.S.) 59 (4), 393–408, 2002.[8] M. I. Gomes, L. de Haan, and D. Pestana, Joint exceedances of the ARCH process,J. Applied Probab. 41 (3), 919–926, 2004.[9] M. I. Gomes, L. de Haan, and D. Pestana, Correction: ‘Joint Exceedances of the ARCHProcess’, J. Applied Probab. 43 (4), 1206, 2006.]]></page><page Index="22" isMAC="true"><![CDATA[12 M. I. GOMES[10] M. I. Gomes, A. Hall, and C. Miranda, Subsampling techniques and the jackknife methodology in the estimation of the extremal index, Comput. Statist. Data Anal. 52 (4), 2022–2041, 2008.[11] M. I. Gomes, M. J. Martins, and M. M. Neves, Generalised Jackknife-based estimators for univariate extreme-value modeling, Comm. Statist. Theory Methods 42 (7), 1227– 1245, 2013.[12] E. Gonc¸alves and N. Mendes-Lopes, Modelos GARCH e TARCH: estacionaridade forte, estacionaridade fraca, ergodicidade e comportamento limite do agregado temporal, Por- tugal. Math. 50 (4), 448–465, 1993.[13] E. Gon¸calves and N. Mendes-Lopes, Distributional Properties of Generalized Threshold ARCH Models, Selected Papers of Statistical Societies, in: Advances in Regression, Sur- vival Analysis, Extreme Values, Markov Processes and Other Statistical Applications, J. Lita da Silva, F. Caeiro, I. Natrio, C. A. Braumann (eds.), Springer-Verlag, Berlin Heidelberg, 213–222, 2013.[14] E. Gon¸calves, J. Leite, and N. Mendes-Lopes, On the finite dimensional laws of threshold GARCH processes, in: Recent Developments in Modeling and Applications in Statistics, P. E. Oliveira, M. G. Temido, C. Henriques, M. Vichi (eds.), Springer-Verlag, Berlin Heidelberg, 237–247, 2013.[15] L. de Haan, S. Resnick, H. Rootz´en, and C. de Vries, Extremal behaviour of solutions to a stochastic di↵erence equation with applications to ARCH-processes, Stoch. Processes Appl. 32, 213–224, 1989.[16] A. O. Hall, Maximum term of a particular autoregressive sequence with discrete margins, Comm. Statist. Theory Methods 25 (4), 721–736, 1996.[17] H. Kesten, Random di↵erence equations and renewal theory for products of random matrices, Acta Math. 131, 207-248, 1973.[18] M. R. Leadbetter, G. Lindgren, and H. Rootz´en, Extremes and Related Properties of Random Sequences and Processes, Springer-Verlag, Berlin, 1983.[19] M. R. Leadbetter and L. Nandagopalan, On exceedance point process for stationary sequences under mild oscillation restrictions, in: Extreme Value Theory, Proceedings, Oberwolfach 1987, J. Hu¨sler, R. D. Reiss (eds.), Lecture Notes in Statist. 52, 69–80, Springer-Verlag, Berlim, 1989.[20] N. Mendes-Lopes, Processus ponctuels chromatiques: estimation de la r´epartition locale des couleurs, Publ. Inst. Statist. Univ. Paris 28 (3), 39–58, 1983.[21] N. Mendes-Lopes, Convergence et optimisation d’un estimateur de la r´epartition locale des couleurs d’un processus ponctuel chromatique, Publ. Inst. Statist. Univ. Paris 29 (2), 49–68, 1984.[22] S. Nandagopalan, Multivariate Extremes and Estimation of the Extremal Index, Ph.D. Thesis, Univ. North Carolina at Chapel Hill, 1990.[23] M. G. Temido, Classes de Leis Limites em Teoria de Valores Extremos – Estabilidade e Semiestabilidade, Tese de Doutoramento, Departamento de Matem´atica, Universidade de Coimbra, 2000.[24] W. Vervaat, On a stochastic di↵erential equation and a representation of nonnegative infinitely divisible random variables, Adv. in Appl. Probab. 11, 750–783, 1979.(M. I. Gomes) DEIO, FCUL, 1749-016, Lisboa, Portugal E-mail address: ivette.gomes@fc.ul.pt]]></page><page Index="23" isMAC="true"><![CDATA[MAX-MIN DEPENDENCE COEFFICIENTS FOR MULTIVARIATE EXTREME VALUE DISTRIBUTIONSHELENA FERREIRADedicated to Nazar´e, my teacher of Probability, matrix of inspiration and energyAbstract. We measure the dependence among subvectors of a random vector with Multivariate Extreme Value distribution by using the ex- pected value of a range and relate this coe cient of dependence with the multivariate tail dependence and extremal coe cients. The introduced coe cient extends the concept of madogram for several locations and several regions. The results are illustrated with some usual distributions and applied to financial data.1. IntroductionThe dependence structure of a Multivariate Extreme Value (MEV) distri- bution is completely characterised by its dependence function (Resnick, 1987; Beirlant et al., 2004). Since this function cannot be easily inferred from data the dependence coe cients are useful, despite the fact that one coe cient cannot preserve all the information about this function.The most popular of the dependence coe cients are those based on the tail dependence (Sybuya, 1960; Li, 2009). They summarize the probability of occurrence of extreme values for one or more random variables given that an- other(s) assumes extreme values too. For the MEV distributions the extremal coe cient (Tiago de Oliveira, 1962-63; Smith, 1990) is certainly a crucial and perhaps insurmountable tool when we have to summarize the dependence. For a d dimensional random vector we have 2d  d extremal coe cients whose con- sistency properties are discussed in Schlather and Tawn (2002). For an overview of other dependence measures see, for instance, Joe (1997).Accepted: 17 February 2015.2010 Mathematics Subject Classification. 60G70.Key words and phrases. Multivariate Extreme Value distribution, dependence, range. The work was supported by PEst-OE/MAT/UI0212/2014.13]]></page><page Index="24" isMAC="true"><![CDATA[14 H. FERREIRATo the best of our knowledge, there is no extremal dependence coe cients for p   3 subvectors X1,...,Xp of a random vector X = (X1,...,Xd) with MEV distribution.The need to evaluate the strength of dependence among subvectors arises forinstanceintheSsettingofmax-stablerandomfields.Let{Xi}i2R2 bea max-stable random field and I1,...,Ip sets of locations in R2. The joint dis- tribution of Xi,i 2 pj=1 Ij, is a MEV distribution and we want to summarize the dependence among the grouped values {Xi, i 2 Ij } at di↵erent regions Ij , j = 1,...,p. This problem is treated by several authors for two variables Xi and Xj corresponding to two locations i and j (see Naveau et al. (2009) and references therein) and the obtained results are extended for two regions I1 and I2 in Fonseca et al. (2015).In finance, we are frequently interested in assessing the dependence among several big world markets, considering each one as a random subvector. For an application with grouped financial stock markets see for instance Ferreira and Ferreira (2012a).We propose to evaluate the degree of dependence among subvectors X1, . . . , Xp of X with MEV distribution by using an expected range, which will be referred as a “max-min coe cient”. It is a summary measure that takes into account the whole group of the extremal coe cients ✏(Xj ) of Xj , j = 1, . . . , p. Our approach is an extension of the modeling for pairwise dependence through- out the madogram (Poncet et al., 2006; Naveau et al., 2009), an extreme-value analogue of the variogram (Cressie, 1993), since it enables to summarize the spatial dependence structure for several locations or regions of locations.The proposed moment-based dependence tool takes into account the spread and dependence among the subvectors and can be easily estimated.The paper is organized as follows. We introduce in Section 2 the dependence coe cient which is well defined for any random vector with MEV distribution, is a function of its copula and is invariant with respect permutations of the variables. Its relations with the multivariate tail dependence and the extremal coe cients are presented. Based on the expected range coe cient considered we compare a MEV distribution with others more concordant distributions and state some bounds.In Section 3, we compute the max-min coe cients for the multivariate mar- ginal distribution of the Multivariate Maxima of Moving Maxima process and the Symmetric Logistic distribution. We refer briefly an estimator for the max- min coe cients and apply it to grouped financial stock markets.]]></page><page Index="25" isMAC="true"><![CDATA[MAX-MIN DEPENDENCE COEFFICIENTS 152. Max-min dependence coefficientsLet X = (X1, . . . , Xd) be a vector of unit Fr´echet random variables, that is, with marginal distribution function F(x) = exp( x 1), x > 0, and G denote the Multivariate Extreme Value distribution of X.The tail dependence function (Huang, 1992; Schmidt and Stadtmu¨ller, 2006) of G is defined by `(x ,...,x ) =  logG(x 1,...,x 1), (x ,...,x ) 2 [0,1)d.1d1d1d It is a convex function, homogeneous of order one and satisfies_d j=1xj  `(x1,...,xd) Xd j=1xj,with the lower bound corresponding to X with totally dependent margins and the upper bound to X with independent margins.The tail dependence function `(x1,...,xd) gives us informaWtion about the probability of occurrence of extreme events for the maximum dj=1 F(Xj). In fact, we have (Schmidt and Stadtmu¨ller, 2006; Ferreira and Ferreira, 2012a, 2012b)`(x1,...,xd)= lim ⇣ tlogP⇣F(X1)1 x1,...,F(Xd)1 xd⌘⌘ t!1⇣ xt x⌘t= lim tP F(X1)>1  1 _···_F(Xd)>1  d . t!1 t tIt holds that `(x,...,x) = x`(1,...,1) and, for  i(S) = 1 if i 2 S and  i(S) = 0 i f i 2/ S ,`( 1(S),..., d(S)) = ✏(XS), (2.1)where ✏(XS) is the extremal coe cient of the subvector XS of X with indicesin S (Tiago de Oliveira, 1962-63; Smith, 1990). It takes values in the interval[1,|S|], where |S| is the number of elementsVin S, with ✏(XS) = 1 when XS hasthe minimum copula CXS (u1,...,ud)S = Q uj and ✏(XS) = |S| when XS j2Shas the product copula CXS (u1,...,ud)S = j2S uj. WLet I = {I1,...,Ip} be a partition of D = {1,...,d}, M(Ij) = i2Ij Xi andXIj the subvector of X with indices in Ij.For given  j 2 (0,1), j = 1,...,p, we will summarize the extremal de-pendence among the weighted subvectors 1 XIj , j = 1, . . . , p, through the  jcoe cient R(X,  , I) defined as follows.Definition 2.1. Let X be a vector of unit Fr´echet random variables and Mul- tivariateExtremeValuedistribution.Foreach =( 1,..., p)2(0,1)p and]]></page><page Index="26" isMAC="true"><![CDATA[16 H. FERREIRA partition I = {I1,...,Ip} of D, we defineR(X,  , I ) = E We remark the relationsF  j (M (Ij ))  ^p j=1F  j (M (Ij )).R(X,  , I ) = E 0@ _{i,j }⇢{1,...,p}0@ _p j=11A F  j (M (Ij ))   F  i (M (Ii )) 1A (2.2)andR(X, ,I)=E0@_p _ F j(Xi)  ^p _ F j(Xi)1A. (2.3)j=1 i2Ij j=1 i2IjBy taking p = 2 in (2.2), we find in 1 R(X,  , I) the generalized madogram2introduced in Fonseca et al. (2015), which in turns is the the  -madogram (Naveuetal.,2009)whend=2=pand1  2 = 1 2(0,1).The max-min coe cient and the generalized madograms for pairs of sets Ii and Ij can be related throughoutR(X, ,I)   _ R((XIi,XIj ),( i, j),{Ii,Ij}). {i,j }⇢{1,...,p}We first present a key result that relates the expectation of Wdj=1 F  j (Xj ) with the tail dependence function of G, which enables the derivation of the main properties of R(X, ,I). The result also points out that in this work we can assume that the MEV distributions have unit Fr´echet margins without loss of generality.Proposition 2.2. Let X be a random vector with MEV distribution G, unit Fr´echet F margins and tail dependence function `. If Y has MEV distribution with marginal distributions Fj, j = 1,...,d, and the same copula as X, then for each   2 (0,1)d, it holds that0 _d 1 0 _d 1 ` (     1 , . . . ,     1 ) E @ F  j (Yj )A = E @ F  j (Xj )A = 1 d. (2.4)j 1+`(  1,...,  1)j=1 j=11 d]]></page><page Index="27" isMAC="true"><![CDATA[MAX-MIN DEPENDENCE COEFFICIENTS 17Proof. We first deduce the distribution of Wd F  j (Yj ). Denoting the copula j=1 jof X by CX, we have, for each u 2 [0,1],P@ j A 1  dFj (Yj )  u = CX u , . . . , u = G     1 log u , . . . ,   1 log u0_d 1⇣ 1  1⌘✓1 1◆j=1 1d = exp  ` ( logu)  1,...,( logu)  1  = u`(  1,...,  1). 1dThenThe next result shows that the max-min coe cient takes into account the taildependencefunctionofallsubvectorsX[j2TIj ofXwithindicesin[j2TIj,;=6 T✓{1,...,p}.From the Proposition 2.2 and (2.1) it holds that, for each ; =6 T ✓ {1, . . . , p},0_d 1Z1( 1 1)    E@ F j(Y )A= u`  1 ,..., d `   1,...,  1 dujj1d j=1 0`(  1,...,  1) =1d.1d1+`(  1,...,  1) 1d0 1 `0@X  1 1(Ij),...,X  1 d(Ij)1A __jj@  j A j2T j2TE F (Xi) = 0X X 1, (2.5)j2T j2Tleading to the following relations of the max-min coe cients with the tail de- pendence and the extremal coe cients. For sake of simplicity we denote the above expectation by e (Ij , j 2 T ).Proposition 2.3. If X has MEV distribution then, for each partition I ={I1,...,Ip} of D and   2 (0,1)p, it holdsXthatR(X,  , I ) = e (Ij , j 2 {1, . . . , p})   ( 1)|T |+1 e (Ij , j 2 T );6=T ✓{1,...,p}(2.6)(2.7)and, for   = 1 = (1,...,1),✏(X) X( 1)|T |+1✏(X[j2T Ij )R(X,1,I)=1+✏(X) ;6=T ✓{1,...,p}1+✏(X[ I ). j2T j⇤j2T i2Ij 1+`@   1 1(Ij),...,   1 d(Ij)A jj]]></page><page Index="28" isMAC="true"><![CDATA[18 H. FERREIRAProof. To obtain the first equality we first apply in (2.3) the relation ^p _ F j (Xi) = X ( 1)|T|+1 _ _ F j (Xi)j=1i2Ij ;=6 T✓{1,...,p} j2T i2Ijand then the Proposition 2.1 with (2.5). The statement in (2.7) is a consequenceof (2.6) and (2.1). ⇤ For the particular case of Ij = {j}, j = 1,...,d, we denote R(X, ,I) simplyby R(X,  ) and we have, as a consequence of the above result, that `(  1,...,  1) X `(  1  (S),...,  1  (S))R(X, )= 1 d   ( 1)|S|+1 1 1 d d 1+`(  1,...,  1) 1+`(  1  (S),...,  1  (S))and1 d;6=S✓{1,...,d} 11 ddR(X,1) = ✏(X)   X ( 1)|S|+1 ✏(XS) . 1 + ✏(X) ;6=S✓{1,...,d} 1 + ✏(XS )These two relations extend the result of the Proposition 1 in Naveau et al. (2009) and equation (14) in Cooley et al. (2006), where, for d = 2, 1    2 =  1 =   2 (0, 1), we have1 `(1,1)31 R(X, )=   1    2 1+`(1, 1 ) 2(1+ )(2  )   1  and R(X, 1) = ✏(X)   1 . ✏(X)+1Our next step is to compare the value of the max-min coe cient R(X, ,I) with the corresponding coe cient in the two boundary cases of independent or totally dependent XIj , j = 1,...,p.Proposition 2.4. Let X, Xˆ = (Xˆ1,...,Xˆd) and X¯ = (X¯1,...,X¯d) be vec- tors of unit Fr´echet random variables with MEV distributions such that Xˆ Ij , j = 1,...,p, are independent, X¯Ij , j = 1,...,p, are totally dependent and, for each j = 1, . . . , p, Xˆ Ij , X¯ Ij and XIj are identically distributed. Then, for each   2 (0, 1)d and partition I of D, it holds that(a) R(X¯, ,I)  R(X, ,I)  R(Xˆ, ,I), ( b ) R ( X¯ , 1 , I ) = 0 ,]]></page><page Index="29" isMAC="true"><![CDATA[^p0@\p 1A Yp {M(Ij) > xj}  MAX-MIN DEPENDENCE COEFFICIENTS 19p✏(XIj ) ✏(XIj )j=1 |T|+1 j2T XX X Xˆ(c) R(X,1,I)=Proof. If X has MEV distribution then it is associated (Marshall and Olkin, 1983) and then the variables M (Ij ), j = 1, . . . , p, are also associated (Esary et al., 1967). For these associated variables it holds thatp   ( 1) 1+X✏(XIj). 1 + ✏(XIj ) ;6=T ✓{1,...,p} j2Tj=1P (M(Ij) > xj)   Pand 0 1P (M(Ij) > xj)j=1j=1 j=1^p \p YpP (M(Ij)  xj)   P @ {M(Ij)  xj}A   P (M(Ij)  xj).j=1 j=1 j=1By taking Mˆ(Ij) = Wi2Ij Xˆi and M¯(Ij) = Wi2Ij X¯i, j = 1,...,p, we can rewrite the above inequalities as0\p 1 0\p 1 0\p 1 P@ {M¯(Ij)>xj}A P@ {M(Ij)>xj}A P@ {Mˆ(Ij)>xj}Aj=1 j=1 j=1and0\p 1 0\p 1 0\p 1 P@ {M¯(Ij)xj}A P@ {M(Ij)xj}A P@ {Mˆ(Ij)xj}A.j=1 j=1 j=1This concordance order implies (Shaked and Shanthikumar, 2007) that0 0 _p E @ f @1 1 0 0 _p M¯ ( I j ) A A  E @ f @1 1 0 0 _p 1 1 M ( I j ) A A  E @ f @ Mˆ ( I j ) A A ,j=1and 0 0^p 11 0 0^p 11 0 0^p 11E @ f @ M¯ ( I j ) A A   E @ f @ M ( I j ) A A   E @ f @ Mˆ ( I j ) A A j=1 j=1 j=1for all non-decreasing functions f. The result in (a) follows by taking f = Fand replacing Xi by Xi , i 2 Ij, for each j = 1,...,p.  jj=1j=1]]></page><page Index="30" isMAC="true"><![CDATA[20 H. FERREIRAThe equalities in (b) and (c) follow from (2.7) and, in particular for (b), we recall that if X¯ Ij , j = 1, . . . , p, are totally dependent vectors then the copula of X¯ is also the copula of the minimum (Nelsen, 2006). ⇤The arguments in the proof of (a), of the previous proposition, can be applied to prove that R(X,  , I) is a dependence coe cient that decreases with respect to the multivariate concordance ordering.For the particular case of Ij = {j},Xj = 1, . . . , d, the equality in (c) leads to R(Xˆ,1)= d   ( 1)|S|+1 |S|d + 1 ;=6 S✓{1,...,d} 1 + |S| dXd   kd 1=d+1  ( 1)k+1 dk k+1=d+1, k=1which extends the already known result for the case of d = 2, where R(Xˆ , 1) = 1.3Therefore, if we define⇢ = 1   d + 1 R ( Xˆ , 1 )d 1we have a dependence coe cient with several useful properties: a) several vari- ables can be taken into account; b) it takes values in [0,1] and higher values indicate stronger dependence; c) it is independent of the univariate marginal distributions; d) it can be related with other coe cients in the literature such as the tail dependence and the extremal coe cients; e) it agrees with the concor- dance property for multivariate distributions; f) it has as a particular case the variogram from geostatistics; g) it can be easily implemented and estimated,as we will see in the next section.3. Examples and an applicationIn order to illustrate the previous results, we consider two families of Mul- tivariate Extreme Value distributions and we compute the expressions for R(X,  , I), which can be easily implemented.Example 3.1. If X is the MEV marginal distribution of the Multivariate Max- ima of Moving Maxima processes considered in Smith and Weissman (1996) then `(x1, . . . , xd) = P1l=1 P1k= 1 Wdj=1 xj ↵l,k,j , where the ↵l,k,j are real non- negative constants that sum one. For this tail dependence function we obtain]]></page><page Index="31" isMAC="true"><![CDATA[and, in particular,Xl=1 k= 1 t2T j2It ( 1)|S|+1  X1 X1 _d 1+R(X, )=;6=S({1,...,d} 1+X1 X1 _ l=1 k= 1 j2S  1 ↵l,k,j jl,k,jR(X, )=01✓   0 d 1✓MAX-MIN DEPENDENCE COEFFICIENTS21l,k,j.X ( 1)|T |+1 R ( X ,   , I ) = X1 X1 _ _;6=T ({1,...,p} 1+   1 ↵ t1 + ( 1)p   X1 X1 _p _X ( 1)|T |+1 1 + ( 1)p;=6 T({1,...,p} 1+   1/✓|I | 1+   1/✓|I | tt jtl,k,j1+l=1 k= 1 t=1j2It 1 + ( 1)d  1 ↵ j  1 ↵ tl=1 k= 1 j=1 Example3.2. FortheSymmetricLogisticmodel`(x1,...,xd)=⇣Pd x1/✓⌘✓we haveR(X, ,I)= X !✓  Xp !✓,j=1 jt2T t=1 X ( 1)|S|+1 1 + ( 1)d;6=S({1,...,d} 1+@X  1/✓A 1+@X  1/✓A jjj2S j=1 Xd 1   1 1andR(X,1)= ( 1)k+1 dk 1+k✓  (1+( 1)d)1+d✓.k=1Several parametric and non-parametric estimators for the tail dependence function are available in the literature (Beirlant et al., 2004; Schmidt and Stadtmu¨ller, 2006; Krajina, 2010) which can be applied to the terms in (2.6). The comparison of estimation procedures is out of our purposes in this paper and we simply remark that the definition of the max-min dependence coe cient suggests a non-parametric estimator based on sample means.Let X(k) = (X(k),...,X(k)), k = 1,...,n, be a sequence of independent 1dcopies of X and Fˆj the empirical distribution provided by X(k), k = 1, . . . , n,j = 1,...,p.j]]></page><page Index="32" isMAC="true"><![CDATA[22 H. FERREIRAA natural estimator for R(X,  , I ) isRˆ ( X ,   , I ) = 1 Xn 0@ _p _ Fˆ   j ( X ( k ) )   ^p _ Fˆ   j ( X ( k ) ) 1A ,niiii k=1 j=1 i2Ij j=1 i2Ijand, in particular for p = d, 01 Rˆ(X,  ) = @ Fˆ j (X(k))   Fˆ j (X(k))A .1WXn _d ^d njjjjj=1If we denote Mˆk(Ij) j = Fˆ j (X(k)) then we can write1 Xn _pRˆ(X, ,I) = Mˆk(Ij) j  X( 1)|T|+11 Xn _ n k=1 j2TMˆk(Ij) j .k=1 j=1 i2Ij i in k=1 j=1;6=T ✓{1,...,p}The strong consistency of the terms of this sum is stated in the proof of the Proposition 3.8 of Ferreira and Ferreira (2012b) and the asymptotic normality can be deduced from the Theorem 6 in Fermanian et al. (2004).As an application of this estimation procedure we consider for X some financial stock markets grouped in the three big world markets I1 =Europe, I2 =USA and I3 =Far East, as considered in Ferreira and Ferreira (2012b). The data are monthly maximums of the negative log-returns of the closing values of the stock market indexes CAC 40 (France), FTSE100 (UK), SMI (Swiss), XDAX (German), Dow Jones (USA), Nasdaq (USA), SP500 (USA), HSI (China) and Nikkei (Japan), from January 1993 to March 2004.In the Table 1 we present the estimates corresponding to M¯(A)=1Xn _Fˆi(X(k))which we need to compute Rˆ(X, 1, I) for I = {I1, I2, I3}. Table 1. I1 =Europe, I2 =USA and I3 =Far East.We obtain then Rˆ(X,1,I) = 0.321, Rˆ((XI1,XI2),1,{I1,I2}) = 0.172, Rˆ((XI1,XI3),1,{I1,I3}) = 0.222 and Rˆ((XI2,XI3),1,{I2,I3}) = 0.247, sug- gesting a stronger dependence between I1 and I2.nj k=1 i2AAI1I2I3I1 [ I2 0.739I1 [ I3 0.770I2 [ I3 0.744I1 [I2 [I3 0.801M¯ ( A )0.6920.6140.626]]></page><page Index="33" isMAC="true"><![CDATA[MAX-MIN DEPENDENCE COEFFICIENTS 23References[1] J. Beirlant, Y. Goegebeur, J. Segers, and J. L. Teugels, Statistics of Extremes: Theory and Applications, England, John Wiley & Sons, 2004.[2] N. A. C. Cressie, Statistics for spatial data, Revised reprint of the 1991 edition, John Wiley & Sons, Inc., New York, 1993.[3] J. D. Esary, F. Proschan, and D. W. Walkup, Association of Random Variables, with Applications, Ann. Math. Statist. 38 (5), 1466–1474, 1967.[4] M. Falk, J. Hu¨sler, and R.-D. Reiss, Laws of Small Numbers: Extremes and Rare Events, 3rd ed., Birkh¨auser, Basel, 2010.[5] J.-D. Fermanian, D. Radulovi´c, and M. Wegkamp, Weak convergence of empirical copula processes, Bernoulli 10 (5), 847–860, 2004.[6] H. Ferreira and M. Ferreira, Fragility Index of block tailed vectors, J. Statist. Plann. Inference 142 (7), 1837–1848, 2012a.[7] H. Ferreira and M. Ferreira, On extremal dependence of block vectors, Kybernetika 48 (5), 988–1006, 2012b.[8] C. Fonseca, L. Pereira, H. Ferreira, and A. P. Martins, Generalized madogram and pairwise dependence of maxima over two regions of a random field, Kybernetika (in press), 2015[9] X. Huang, Statistics of Bivariate Extreme Values, Ph.D. thesis, Tinbergen Institute Research Series 22, Erasmus University Rotterdam, 1992.[10] H. Joe, Multivariate Models and Dependence Concepts, Chapman & Hall//CRC, Lon- don, 1997.[11] A. Krajina, An M-Estimator of Multivariate Tail Dependence, Ph.D. thesis, Tilburg: Tilburg University Press, 2010.[12] H. Li, Orthant tail dependence of multivariate extreme value distributions, J. Multivari- ate Anal. 100 (1), 243–256, 2009.[13] A. W. Marshall and I. Olkin, Domains of attraction of multivariate extreme value dis- tributions, Ann. Probab. 11 (1), 168–177, 1983.[14] P. Naveau, A. Guillou, D. Cooley, and J. Diebolt, Modelling pairwise dependence of maxima in space, Biometrika 96 (1), 1–17, 2009.[15] R. B. Nelsen, An Introduction to Copulas, Second Edition, Springer, New York, 2006.[16] P. Poncet, D. Cooley, and P. Naveau, Variograms for spatial max-stable random fields, in: Dependence in probability and statistics, P. Bertail, P. Doukhan, P. Soulier (eds.),Lecture Notes in Statist. 187, 373–390, Springer, New York, 2006.[17] S. Resnick, Extreme Values, Regular Variation and Point Processes, Springer-Verlag,New York, 1987.[18] M. Schlather and J. A.Tawn, Inequalities for the extremal coe cients of multivariateextreme value distributions, Extremes 5 (1), 87–102, 2002.[19] M. Schlather and J. A. Tawn, A dependence measure for multivariate and spatialextreme values: Properties and inference, Biometrika 90, 139–156, 2003.[20] R. Schmidt and U. Stadtmu¨ller, Nonparametric estimation of tail dependence, Scand.J. Stat. 33, 307–335, 2006.[21] M. Shaked and J. G. Shanthikumar, Stochastic orders, Springer-Verlag, New York, 2007.[22] M. Sibuya, Bivariate extreme statistics, Ann. Inst. Statist. Math. 11, 195–210, 1960.[23] R. L. Smith, Max-stable processes and spatial extremes, Preprint, Univ. North Carolina,USA, 1990.]]></page><page Index="34" isMAC="true"><![CDATA[24 H. FERREIRA[24] R. L. Smith and I. Weissman, Characterization and estimation of the multivariate ex- tremal index, Technical Report, Univ. North, Carolina, 1996.[25] J. Tiago de Oliveira, Structure theory of bivariate extremes, extensions, Estudos de Math. Estat. Econom. 7, 165–195, 1962/63.(H. Ferreira) Department of Mathematics, University of Beira Interior, Covilha˜, PortugalE-mail address: helenaf@ubi.pt]]></page><page Index="35" isMAC="true"><![CDATA[THE TAYLOR PROPERTY IN NON-NEGATIVE AUTOREGRESSIVE AND BILINEAR STOCHASTIC PROCESSESESMERALDA GONC¸ALVES AND CRISTINA M. MARTINSDedicated to our dear friend Nazar´eAbstract. Inthiswork,wecomebacktotheanalysisoftheTaylorprop- erty in linear and bilinear models. Limiting the study to non-negative models, we begin by recalling the main theoretical results on its occur- rence in autoregressive and diagonal bilinear models of order one. These results are discussed in models with significantly di↵erent values of the kurtosis coe cient. In both classes, it is possible to relate the presence of the Taylor property and the kurtosis of the model. Moreover, in the autoregressive case the property is present if and only if the generator process is symmetric or right-skewed distributed.1. IntroductionTaylor e↵ect is a characteristic present in temporal series of diverse nature.This stylized fact was, for the first time, detected by Taylor (1986) in the returnsof some financial series. Taylor found that the autocorrelations of the absolutereturns were systematically higher than the corresponding autocorrelations ofsquared observations. Namely, considering T observations, X1 , X2 , ..., XT , fromaprocessX=(X,t2Z),Taylorfoundthat ⇢b (h)>⇢b 2 (h),h=1,2,..., t |X|Xwhere ⇢bX denotes the empirical autocorrelation function of X. This fact is now known as Taylor e↵ect.The corresponding theoretical relation is called the Taylor property, and assuming that the functions ⇢|X| and ⇢X2 are positive, is expressed by theAccepted: 23 March 2015.2010 Mathematics Subject Classification. 62M10.Key words and phrases. Autoregressive and Bilinear models, kurtosis, Taylor property.This work was partially supported by the Centre for Mathematics of the Univer- sity of Coimbra – UID/MAT/00324/2013, funded by the Portuguese Government through FCT/MEC and co-funded by the European Regional Development Fund through the Part- nership Agreement PT2020.25]]></page><page Index="36" isMAC="true"><![CDATA[26 E. GONC¸ALVES AND C. M. MARTINScondition⇢|X|(h)>⇢X2 (h), h2N,with ⇢X denoting the autocorrelation function of X.The presence of this property in a certain class of models for time series canhelp in the selection of a more appropriate model to describe the dynamics of the series of interest. However, this study requires the knowledge of moments of X of order higher than 2 which makes such analysis theoretically di cult.Since the Taylor e↵ect was initially detected in financial time series, it is nat- ural that the first theoretical studies have focused on models used in the analysis of these series, specifically the conditionally heteroskedastic models, as can be found in He and Ter¨asvirta (1999), Gon¸calves, Leite and Mendes-Lopes (2009), Haas (2009) and Leite (2013), whose studies involve the class of Generalized Threshold Autoregressive Conditionally Heteroskedastic (GTARCH) models.Interestingly, these studies have shown that the presence of the Taylor prop- erty appears to have a stronger relationship with the weight of the tails of the process than with the characteristic of heteroskedasticity. This finding leads us to question the presence of the Taylor property in models without conditional heteroskedasticity but still suitable for financial time series analysis.The modeling of non-negative time series has been widely regarded in the literature (Zang and Tong, 2001) and in Tsay and Chan (2007) the interest of non-negative ARMA processes in the context of stochastic modeling of fi- nancial time series is referred. Additionally, the recent finding that the Taylor e↵ect is clearly present in some non-negative physical time series (Gonc¸alves et al., 2014) also leads us to question the presence of the property in bilinear models, which reveal themselves useful in the analysis of time series recorded, for instance, in seismology, hydrology and astronomy.So, this paper focuses on non-negative stochastic models and, given the complexity of the calculations, we limit our study to the first-order AR and bilinear models.Returning to Gonc¸alves, Martins and Mendes-Lopes (2014, 2015) we begin by recalling the fundamental results on the occurrence of the Taylor property in non-negative autoregressive and bilinear models of order 1. We point out that the property is analyzed for all lag h in the case of linear models. On the other hand, in the case of the bilinear model our study is limited to h = 1, due to the di culty associated with the calculations of the autocorrelations. The presence of the Taylor property is then assured in the linear case when the error process is symmetrically or right-skewed distributed and new examples are developed considering autoregressive and bilinear models with error processes presenting di↵erent values of kurtosis, being notorious the importance of this parameter on the presence of the Taylor property in the model.]]></page><page Index="37" isMAC="true"><![CDATA[THE TAYLOR PROPERTY IN AR AND BILINEAR PROCESSES 27In what follows, we denote by KX the kurtosis coe cient of Xt, that is, KX = μ4,X/(μ2,X)2, where μk,X is the centered moment of order k of Xt.2. The Taylor property in first-order non-negative autoregressive modelsLet X = (Xt, t 2 Z) be a real stochastic process such thatXt =  Xt 1 + "t (2.1)where 0 <   < 1 and " = ("t , t 2 Z) is a sequence of non-negative identically distributed random variables with moments up to the fourth order and such that E  "it|"t 1  = mi, i = 1, 2, 3, t 2 Z, with "t 1 the  -algebra of the past of "t, that is, "t 1 =   ("t 1, "t 2, ...). So, E  "it  = E  E  "it|"t 1   = mi, i = 1, 2, 3. Naturally, let us denote E  "4t   by m4, t 2 Z.The following lemma summarizes, in these conditions, the expressions of the autocorrelation functions of X and X2 for h 2 N.Lemma 2.1. (Gon¸calves, Martins and Mendes-Lopes, 2014) We havea) ⇢X (h)= h, h2N.2 h   h     2 h C o v X t , X t2+ 2 m 1 1     V ( X t2 ) , h 2 N .  b ) ⇢ X 2 ( h ) =  We note that ⇢X2 is positive, since Cov  Xt, Xt2    0 as Xt is non-negative.The following result follows immediately.Theorem 2.2. (Gon¸calves, Martins and Mendes-Lopes, 2014) The process X defined in (2.1) satisfies the Taylor property if and only ifCov Xt,Xt2  < 1  . V ( X t2 ) 2 m 1The studies on the Taylor property previously referred have shown a strong relation between its occurrence and strong values of the kurtosis of the process X. In order to explore this relation, we rewrite the previous condition in terms of this coe cient.Corollary 2.3. The process X defined in (2.1) satisfies the Taylor property if and only ifKX >1 2E(Xt)μ3,X. V ( X t2 )Let us note that, as KX   1, we may conclude that if μ3,X   0 then the Taylor property is present in the process X, or equivalently, if the genera- tor process is symmetric or right-skewed distributed, taking into account that μ3,X = μ3,"/(1    3).]]></page><page Index="38" isMAC="true"><![CDATA[28 E. GONC¸ALVES AND C. M. MARTINSWe also note that the kurtosis of the processes X and " are related by 6 2 1  2So, the process X is leptokurtic if the generator process, ", is leptokurtic. Moreover, this equality shows that KX is an increasing function of   if K" < 3, and a decreasing function if K" > 3. Also, when   tends to 1, KX tends to 3, independently of the value of K".ExamplesWe evaluate now the presence of the Taylor property in model (2.1) con- sidering generator processes with non-negative distributions and weight tails significantly di↵erent. We take into account the necessary and su cient condi- tion of the presence of the Taylor property, written in the formT ( )=1   2m Cov Xt,Xt2  >0, (2.2) L 1 V ( X t2 )where L denotes the marginal law of the generator process, ". Generator process with platykurtic distributionWe consider the process X defined in (2.1) in cases where "t is distributedaccording to a platykurtic Beta probability law, Beta(↵, ✓), whose density hasthe form f1(x) =  (↵+✓) x↵ 1(1 x)✓ 1 I]0,1[(x), ↵ > 0, ✓ > 0. More precisely,  (↵) (✓)KX( )=1+ 2 +1+ 2K".we analyze the cases ✓ = ↵ (symmetrical distribution), ✓ = 2↵ (right-skewed distribution) and ✓ = ↵ (left-skewed distribution). Concerning the values of2thekurtosis,wehaveK" =31+2↵ when✓=↵,K" =32+3↵ when✓= ↵ and3+2↵ 4+3↵ 2K" =31+3↵ when✓=2↵.Inthecases✓=↵and✓=2↵,Corollary2.32+3↵allows us to conclude that the Taylor property is present. In both cases, thefunctions KX( ) and TL( ) are also functions of ↵ and the presence of the Taylor property becomes, in general, weaker when KX increases.Comparing now the cases ✓ = ↵ and ✓ = 2↵, it is easy to verify that, for each ↵, the kurtosis is greater for the right-skewed distribution.In Figure 1 we present the graphs of the functions KX( ) related to the cases Beta(↵,↵) and Beta(↵,2↵), as well as the graphs of TBeta(↵,↵)( ) and TBeta(↵,2↵)( ), 0< <1, 0<↵<5.]]></page><page Index="39" isMAC="true"><![CDATA[THE TAYLOR PROPERTY IN AR AND BILINEAR PROCESSES 290 3K0 0.3TΑ55Α01 00ΦΦ11Figure 1. Graphs for K = KX( ) (on the left) and T = TL( ) (on the right) in the cases Beta(↵,2↵) (upper surface) and Beta(↵, ↵) (lower surface), 0 <   < 1, 0 < ↵ < 5.This figure shows that the presence of the Taylor property is stronger in the case Beta(↵,2↵), corresponding to greater kurtosis of the model. In the symmetrical case, the presence of the Taylor property is very weak (the values of TBeta(↵,↵)( ) are close to zero).In the case where "t is left-skewed distributed, we have TBeta(↵, ↵ ) ( ) < 0, 2so the Taylor property is not present in the model. In Figure 2 we present the graphforTBeta(↵,↵)( ), 0< <1, 0<↵<5.20 0T 0.3 0ΑΦFigure2. GraphforT=TBeta(↵,↵)( ),0< <1,0<↵<5. 251]]></page><page Index="40" isMAC="true"><![CDATA[30 E. GONC¸ALVES AND C. M. MARTINSGenerator process with leptokurtic distributionWe consider now the process X defined in (2.1) with "t following the Gammadistribution with parameters ✓ and ↵,  (✓,↵), ↵ > 0, ✓ > 0, that is, withdensity f2(x) = ↵✓ e ↵xx✓ 1I]0,+1[ (x), whose kurtosis coe cient is K" =  (✓)3+6.So,K ( )= 6 2 +1  2  3+6 alsodependson✓(butnoton↵)asit ✓ X 1+ 2 1+ 2 ✓happens with T (✓,↵)( ). The Taylor property is always present in model (2.1) as the distribution   (✓, ↵) is right-skewed.We note that when ✓ increases then T (✓,↵) ( ) approaches zero, that is, theautocorrelation functions in study are closer as ✓ increases. This result is inagreement with the evolution of the kurtosis, taking into account that nowthe process is more leptokurtic when ✓ decreases. Analogous conclusions areobtained when we consider "t following a Pareto distribution with parameters↵and✓,Par(↵,✓),withdensityf (x)= ✓↵✓ I (x),↵>0,✓>4(inorder 3 x✓+1 ]↵,+1[to assure the existence of m4). The kurtosis is K" = 3 + 6(✓3+✓2 6✓ 2) which ✓(✓ 3)(✓ 4)is a decreasing function of ✓ that goes to 9 when ✓ tends to infinity, and to infinity when ✓ tends to 4. So, the Pareto distribution is leptokurtic, no matter the value of ✓.These conclusions are illustrated in Figure 3, which shows the graphs of the functions KX ( ), related to the Gamma and Pareto distributions, and those of the functions T (✓,↵) ( ) and TPar(↵,✓) ( ), 0 <   < 1, 4 < ✓ < 10.4 800K10Θ0 04 1T10Θ0 0ΦΦ11Figure 3. Graphs for K = KX( ) (on the left) and T = TL( ) (on the right) in the cases Par(↵,✓) (upper surface) and Gamma(✓,↵) (lower surface), 0 <   < 1, 4 < ✓ < 10.]]></page><page Index="41" isMAC="true"><![CDATA[THE TAYLOR PROPERTY IN AR AND BILINEAR PROCESSES 31We point out that this figure stresses the fact that the presence of the Taylor property is stronger in the Pareto case, that is, its occurrence is more evident when the kurtosis of the model increases.3. The Taylor property in first-order non-negative bilinear modelsWe consider the simple bilinear diagonal modelXt =  Xt 1"t 1 + "t, (3.1)where   > 0 is a real parameter and " = ("t , t 2 Z) a sequence of non-negative i.i.d. random variables.In order to assure the strict and weak stationarity of the processes X = (Xt, t 2 Z) and X2 = (Xt2, t 2 Z), we suppose that E(ln |"t|) and m8 exist and that  4 m4 < 1 , with mi = E("it), i 2 N (Gon¸calves, Martins and Mendes- Lopes, 2015).The nth moment of Xt, nX 4, can be expressed as n ✓n◆whereE(Xn"n) = 1 Xn ✓n◆ n i m E(Xn i"n i), tt 1  nmni=1i n+i t tE(Xn) =  n i m E(Xn i"n i), tiitti=0n  4.It is easy to verify that E(Xt"t) = m2/(1  m1). The values E(Xtn"nt ), n = 1, 2, 3, are obtained recursively by using the previous equation; and finally, we getE(Xtn),forn4.Wenotethat 4m4 <1implies| nmn|<1,n=1,2,3, by Schwarz’s inequality.In this context, the Taylor property is only studied for h = 1, that is, by analyzing if ⇢X (1) > ⇢X2 (1), where ⇢X (1) and ⇢X2 (1) denote, respectively, the autocorrelations of lag 1 of the processes X and X2. It is enough to evaluate E(XtXt 1) and E(Xt2Xt2 1) in order to obtain these autocorrelations. Using (3.1) and the stationarity of the involved processes, we haveE(XtXt 1) =  E(Xt2"t) + E(Xt 1"t)=  E( 2X2 "2 " +2 X " "2 +"3)+E(X " ).t 1 t 1 t t 1 t 1 t t t 1 t Taking into account the independence of the random variables "t, t 2 Z,and the strict stationarity of the related processes, we have E(X2 "2 " ) = 2 2 2 t 1 t 1 tm1E(Xt "t ) and E(Xt 1"t 1"t ) = m2E(Xt"t). ThenE(XtXt 1) =  3m1E(Xt2"2t ) + 2 2m2E(Xt"t) + m1E(Xt) +  m3.]]></page><page Index="42" isMAC="true"><![CDATA[32 E. GONC¸ALVES AND C. M. MARTINSUsing an analogous procedure, we obtainE(X2X2 ) =  4E +2 3E +2 3m E +4 2m E + 2E +2 m E +andE6 =tt 1 1 2 13 14 5 16 + 2m2E(Xt2"2t ) + 2 m1m2E(Xt"t) + m2,whereE1 =E2 =E3 =E4 =E5 =ments of "t. ExamplesIn the following lines, we investigate the presence of the Taylor property in model (3.1), considering some non-negative distributions for the generator process, namely, the uniform distribution in ]0, ↵[, the exponential distribution in ]0,+1[ with mean ↵ and the Pareto distribution, Par(↵,✓), ✓ > 8. In all cases, ↵ is a non-negative parameter and the condition E(|ln"t|) < +1 is satisfied.As in the case of the AR(1) model, the choice of these distributions takes into account the fact that the Taylor property seems to be related with the kurtosis value of the process. The uniform distribution is platykurtic with a constant kurtosis value equal to 1.8, while the exponential distribution is leptokurtic with a constant kurtosis value equal to 9. As we referred before, the kurtosis of the Pareto distribution is greater than 9.We also point out that, in all cases, the condition  4 m4 < 1 and the values of ⇢X(1) and ⇢X2(1) can be written in terms of r = ↵ .In each case, the value of the kurtosis of the process X given by (3.1) also depends on r = ↵ . We present the corresponding graphical representation. We point out that, in all these models, the leptokurtosis of the generator process implies the same property for the process X. In what concerns the Taylor property and kurtosis of X, comparisons are made separately between the first two distributions, uniform and exponential, and also between the ParetoE(Xt2Xt2 1"2t "2t 1)E(Xt2Xt 1"3t "t 1)E(XtXt2 1"t"2t 1)=  2m2E(Xt4"4t ) + 2 m3E(Xt3"3t ) + μ4E(Xt2"2t ),=  2m3E(Xt3"3t ) + 2 m4E(Xt2"2t ) + m5E(Xt"t), =  m1E(Xt3"3t ) + m2E(Xt2"2t ),=  m2E(Xt2"2t ) + m3E(Xt"t),E(XtXt 1"2t "t 1)E(Xt2"4t ) =  2m4E(Xt2"2t ) + 2 m5E(Xt"t) + m6,E(Xt"3t ) =  m3E(Xt"t) + m4.Finally, the values of E(XtXt 1) and E(Xt2Xt2 1) appear in terms of the mo-]]></page><page Index="43" isMAC="true"><![CDATA[THE TAYLOR PROPERTY IN AR AND BILINEAR PROCESSES 33distributions, according to the values of the kurtosis of the process X, that is also, in this last case, a function of ✓.Generator process with uniform distribution in ]0,↵[Inthiscase,thecondition 4m4 <1isequivalentto0<r<p4 5'1.495. For these values of r, the graphs of the di↵erence ⇢X(1) ⇢X2(1) and of KX(r) are presented in Figure 4.15100.040.02- 0.020.2 0.4 0.6 0.81.0 1.2 1.40.2 0.40.6 0.81.0 1.2 1.4Figure 4. Graphs for KX (r) (on the left) and ⇢X (1)   ⇢ 2 (1)p (ontheright)inthecase"t ⇠U(]0,↵[),0<r< 4 5.XFrom Figure 4 (on the right), we can see that the Taylor property is present for values of r in the interval ]1.1868987, p4 5[. So, for a fixed ↵, the Taylor property is achieved for parameterizations of model (3.1) such that  2 #1.1868987, p4 5". ↵↵From Figure 4 (on the left), we observe that the kurtosis of this model is an increasing function of r and that the model is leptokurtic for r > 0.8 (approx.). We also observe that the Taylor property occurs for large values of the kurtosis, namely for KX (r) > 7.403 (' KX (1.1868987)).Generator process with exponential distribution with mean ↵ 41Thecondition  m4 <1isnowequivalentto0<r< p4 24 '0.4518.In this case, model (3.1) presents the Taylor property for parameterizationssuch that  2  0, 0.0695566 [  0.1437879, p1 . ↵ ↵424↵]]></page><page Index="44" isMAC="true"><![CDATA[34 E. GONC¸ALVES AND C. M. MARTINSThis conclusion is illustrated in Figure 5 (on the right). In Figure 5 (on the left), we have the graphical representation of the kurtosis of model (3.1) with exponential generator process.250200 0.20150 100 500.15 0.10 0.050.1 0.20.3 0.40.1 0.20.3 0.4Figure 5. Graphs for KX (r) (on the left) and ⇢X (1)   ⇢X2 (1) 1(ontheright)inthecase"t ⇠E(1/↵),0<r< p424.As in the previous case, the kurtosis of model (3.1) is an increasing function of r but the process X is always leptokurtic in this case. Again, we observe that large kurtosis values correspond to large values of the di↵erence ⇢X (1) ⇢X2 (1). In fact, the Taylor property is clearly present in this model for kurtosis values greater than 16 (' KX (0.16)).We also observe that the kurtosis of the process X is larger when the gener- ator process is exponentially distributed than when it is uniformly distributed, corresponding to an analogous relation between the kurtosis of those generator processes. The Taylor property seems to emerge in a relatively stronger way when the kurtosis of X increases.Generator process with Pareto distributionThe region oqf existence of the autocorrelations in terms of r = ↵  is now defined by0<r<4 ✓ 4,✓>8.✓In Figure 6, we present graphical representations oqf ⇢X (1)   ⇢X 2 (1) andq4 ✓ 4KX (r) for 0 < r < 0.85 and 9  ✓  50. We note that ✓ is an increasingfunctionof✓whoseminimumvaluefor9✓50is 4 5 '0.863. 4]]></page><page Index="45" isMAC="true"><![CDATA[0 400 0.06DK10 9THE TAYLOR PROPERTY IN AR AND BILINEAR PROCESSES 350.8r0.8 r0 9ΘΘ5050Figure 6. Graphs for K = KX(r) (on the left) and D=⇢X(1) ⇢X2(1) (on the right) in the case "t ⇠ Par(↵,✓), 0 < r < 0.85, 9  ✓  50.As can be seen in this figure, the Taylor property is now achieved for all considered parameterizations of model (3.1) and the process X is always lep- tokurtic since KX(r) > 3. These graphical representations also suggest that the presence of the Taylor property is stronger for higher values of the kurtosis of the process X. In fact, as function of ✓, the di↵erence ⇢X (1)   ⇢X2 (1) seems to increase when KX(r) increases, for all values of r that satisfy the condition  4m4 < 1. This situation strongly contributes to conjecture that the Taylor property and leptokurtosis are highly related in time series.4. ConclusionThe studies here presented confirm that linear and bilinear models may mimic the Taylor e↵ect. We note that in the linear case we were able to study the subject in a more complete way than in the bilinear class. In fact, a nec- essary and su cient condition assuring the presence of the Taylor property in the linear class for all lag h in the functions ⇢|X| and ⇢X2 was derived and we were able to conclude its presence in all models with symmetric or right-skewed distributed generator processes.New examples were studied and we still point out the strong relation between the Taylor property and the kurtosis of the process. In fact, in the platykurtic cases studied, the property is not present or is present in a very soft way, becoming more visible as the kurtosis goes away from the reference value of 3. In these cases, the presence of the property is not significant, in the sense]]></page><page Index="46" isMAC="true"><![CDATA[36 E. GONC¸ALVES AND C. M. MARTINSthat the functions ⇢|X| and ⇢X2 are almost equal. Otherwise, in the leptokurtic processes, the Taylor property is present in a significant way, becoming stronger when the kurtosis of the model is also stronger.Again we conclude that the Taylor property is highly dependent on the greater or lesser weight of the tails of the law of the process under study.AcknowledgmentsThe authors are grateful to the referee for his valuable comments.References[1] E. Gonc¸alves, J. Leite, and N. Mendes-Lopes, A mathematical approach to detect the Taylor property in TARCH processes, Statist. Probab. Lett. 79, 602–610, 2009.[2] E. Gonc¸alves, C. M. Martins, and N. Mendes-Lopes, Propriedade de Taylor em processos autorregressivos, Estat´ıstica: A ciˆencia da incerteza, Atas do XXI Congresso Anual da Sociedade Portuguesa de Estat´ıstica, I. Pereira, A. Freitas, M. Scotto, M. E. Silva, C. D. Paulino (eds.), Edi¸c˜oes SPE, 51–53, 2014.[3] E. Gon¸calves, C. M. Martins, and N. Mendes-Lopes, The Taylor property in bilinear models, REVSTAT (in press), 2015.[4] E. Gon¸calves, N. Mendes-Lopes, I. Dorotoviˇc, J. M. Fernandes, and A. Garcia, North and South Hemispheric Solar Activity for Cycles 21-23: Asymmetry and Conditional Volatility of Plage Regions Areas, Sol. Phys. 289 (6), 2283–2296, 2014.[5] J. Leite, Processos Threshold GARCH com potˆencia: estrutura probabilista e aplicac¸o˜es a cartas de controlo, Tese de Doutoramento, Universidade de Coimbra, 2013.[6] M. Haas, Persistence in volatility, conditional kurtosis, and the Taylor property in ab- solute value GARCH processes, Statist. Probab. Lett. 79 (5), 1674–1683, 2009.[7] C. He and H. Tera¨svirta, Properties of moments of a family of GARCH processes, J. Econometrics 92, 173–192, 1999.[8] S. J. Taylor, Modelling Financial Time Series, Jonh Wiley & Sons, Chichester, UK, 1986.[9] H. Tsay and K.S. Chan, A Note on Non-Negative Arma Processes, J. Time Series Anal. 28, 3, 350–360, 2007.[10] Z. Zhang and H. Tong, On some distributional properties of a first order non-negative bilinear time series model, J. Appl. Probab. 38 (3), 659–671, 2001.(E. Gon¸calves) CMUC, Department of Mathematics, University of Coimbra, Por- tugalE-mail address: esmerald@mat.uc.pt(C. M. Martins) Department of Mathematics, University of Coimbra, Portugal E-mail address: cmtm@mat.uc.pt]]></page><page Index="47" isMAC="true"><![CDATA[ON THE ABILITY OF THE POWER THRESHOLD GARCH MODEL TO CAPTURE THE TAYLOR EFFECTJOANA LEITEDedicated to Professor Nazar´e Mendes Lopes on the occasion of her 60th birthdayAbstract. In this paper, we evaluate the capacity that the first order power threshold GARCH ( -TGARCH(1,1)) model has to reproduce the Taylor e↵ect. The Taylor e↵ect is a stylized fact of financial return series which implies that autocorrelations of absolute returns are greater than the ones of squared returns. We firstly consider   = 1 and establish the presence of the Taylor property, Taylor e↵ect theoretical counterpart, for a set of model parameters imposing weak conditions on the generating process distribution. We further explore this set considering particular generating process marginal distributions with di↵erent kurtosis, illus- trating that high values of kurtosis seem to favour the appearance of this property. Finally, considering several values for  , we present an ex- ploratory simulation study which shows that the models incorporated in the  -TGARCH class are not equally favourable to the Taylor e↵ect appearance.1. IntroductionEmpirical regularities shared by a certain time series group are named styl- ized facts and can be used to reveal both strengths and weaknesses of the models proposed for them, as Engle so well summarized in [3].A class of time series that has been extensively analysed in the pursuit of stylized facts are financial daily return series. Some of the stylized facts un- covered are very well known, like (see, for example, Francq and Zakoian [4] orAccepted: 17 February 2015.2010 Mathematics Subject Classification. 62M10, 62P20.Key words and phrases. Financial time series, autocorrelations, Taylor e↵ect, Taylor pro-perty, power TGARCH model.This work was partially supported by the Centre for Mathematics of the Univer-sity of Coimbra – UID/MAT/00324/2013, funded by the Portuguese Government through FCT/MEC and co-funded by the European Regional Development Fund through the Part- nership Agreement PT2020.37]]></page><page Index="48" isMAC="true"><![CDATA[38 J. LEITETaylor [20]): (i) absence of autocorrelations in the returns; (ii) positive corre- lation in squared returns and absolute returns; (iii) volatility clustering, since large absolute returns tend to appear in clusters; (iv) leverage e↵ects, meaning that there is an asymmetric impact of past positive and negative values on the current volatility; and (v) heavy-tailed distribution of returns.The study of stylized facts associated to these series dates back, at least, to 1963, when Mandelbrot [16] identified the volatility clusters. The leverage e↵ects were noted by Black [1], in 1976. These better known stylized facts have very intuitive explanations, as, for example, in the leverage e↵ects case (see [4]), we all now recognize that negative returns (corresponding to price decreases) tend to increase volatility by a larger amount than positive returns (price increases) of the same magnitude.detected by Taylor in 1986. More precisely, while studying 40 return series, Taylor [19] observes that, for n = 1, 2, ..., 30,A relatively more recent and less intuitive stylized fact, that seems to emerge from the combination of volatility clusters and high kurtosis, is the Taylor ef- fect, which involves the sample autocorrelations of power-transformed absolutereturns, ⇢ˆ (k) = cdorr ⇣|r |k , |r |k ⌘ with k > 0, so-called because it was first n tt n⇢ˆn (1) > ⇢ˆn (2) . (1.1)Ding, Granger and Engle [2], Granger and Ding [10], Granger, Spear and Ding [11] and Taylor [20] extend this initial analysis, finding that almost always, for n = 1, 2, ..., 625,⇢ˆn (1) > ⇢ˆn (k) , for some values of k between 0 and 3. (1.2) So Granger and Ding [10] define the Taylor e↵ect to be present in a series if⇢ˆn(1)>⇢ˆn(k), forallnandk6=1. (1.3)Granger [9] adds this to the list of stylized facts that characterize return se- ries dynamics. More studies with further discussion and evidence are cited in Haas [12]. Additionally, we refer that this e↵ect, more specifically the relation (1.1), has also been detect by Gon¸calves, Mendes-Lopes, Dorotoviˇc, Fernandes and Garcia [8] in physical phenomena time series.Taylor e↵ect strong empirical evidence raises the question whether or not empirically relevant volatility models, like the ones in the generalized autore- gressive conditional heteroskedastic (GARCH) class, can capture this e↵ect. In literature, two ways have been considered to address this problem: via simu- lation or using a theoretical approach. Simulations have the advantage of not requiring ⇢n (k) theoretical expressions for each model, however only a small number of model parameterizations can be considered. Since, in the GARCH]]></page><page Index="49" isMAC="true"><![CDATA[ -TGARCH MODEL: CAPTURING THE TAYLOR EFFECT 39class, these theoretical expressions are very hard to obtain and are not available for all k but are indispensable for the latter approach, He and Ter¨asvirta [13] are the first to suggest the analysis of the theoretical relation corresponding to Taylor’s inicial findings, i.e.,⇢n (1) > ⇢n (2), for all n, (1.4)which they named Taylor property.He and Ter¨asvirta [13] consider the conditionally Gaussian first-order ab-solute value GARCH (AVGARCH) model and examine the Taylor property restricted to the autocorrelations of lag n = 1. They find the property to be present for some parameterizations associated with very high values of the model kurtosis. Gon¸calves, Leite and Mendes-Lopes [5] also evaluate this re- stricted version of the property, establishing its presence for some parameter- izations of the first-order threshold ARCH (TARCH) model but without pre- defining the generating process distribution. They also point to the influence the kurtosis of the generating process margins has on the parameterizations set which verify the property, since larger values seem to expand this set. Later, Haas [12] turns back to the first-order AVGARCH model, also not pre-defining the generating process distribution, and fully explores the Taylor property (for any lag n), extending the previous findings.The Taylor property has also been investigated with similar conclusions in stochastic volatility models, by Mora-Gal´an, P´erez and Ruiz [17], Veiga [21], P´erez, Ruiz and Veiga [18] and Malmsten and Ter¨asvirta [15], and also in bilinear models, by Gon¸calves, Martins and Mendes-Lopes [7].In this paper, we consider the first-order power threshold GARCH ( -TGARCH) model, a general model that incorporates many popular GARCH models including the ones already mentioned, and discuss its ability to capture the Taylor e↵ect. Section 2 is dedicated to the presentation of the  -TGARCH model. In Section 3, setting   = 1, the presence of the Taylor property (for any lag) is established for parameterizations of the 1-TGARCH model that can reproduce the leverage e↵ects. Additionally, taking Student’s t-distribution with several degrees of freedom for the variables of the generating process, the influence of its kurtosis in this property is illustrated. Finally, in Section 4, an exploratory simulation study compares  -TGARCH models regarding their ability to reproduce the Taylor e↵ect, considering several values of   and Gauss- ian and Laplace distributions for the generating process margins.2. The first order power threshold GARCH modelLet X = (Xt,t 2 Z) and Z = (Zt,t 2 Z) be real stochastic processes. In the following, Xt+ = max {Xt, 0}, Xt  = max { Xt, 0} and Xt 1 is the  -field]]></page><page Index="50" isMAC="true"><![CDATA[40 J. LEITEgenerated by Xt 1,Xt 2,.... In addition, we consider that the process Z is a sequence of independent and identically distributed random variables with zero mean and unit variance.The process X follows a first-order power   threshold GARCH model, de-noted  -TGAR⇢CH(1,1), if for every t 2 Z, Xt = Zt t          , (2.1)   =↵+↵ X+ +  X  +   t 0 1 t 1 1 t 1 1t 1 where >0,↵0 >0,↵1  0, 1  0, 1  0andZt isindependentofthe Xt 1. Z is called the generating process of X. If  1 = 0, we say that X follows a  -TARCH(1) model.Regarding this specification of  t, first proposed by Ding, Granger and En- gle [2], it is relevant to notice that it allows the model to capture the leverage e↵ect, being enough that ↵1 6=  1, and is more flexible since it does not estab- lish a priori the power   value. Thus, it incorporates, for example, the spec- ificationsofGARCH( =2and↵1 = 1),AVGARCH( =1and↵1= 1), TGARCH (  = 1) and power-GARCH (↵1 =  1) models. In Gon¸calves, Leite and Mendes-Lopes [6] a detailed probabilistic analysis is presented, which in- cludes strict stationarity and stationarity up to  -order of a more general version of this model, namely considering the  -TGARCH(p,q) model with   6= 0 and assuming even less restrictive conditions for the generating process.As previously mentioned, the analysis of the Taylor property requires both ⇢n (1) = corr (|Xt| , |Xt n|) and ⇢n (2) = corr ⇣|Xt|2 , |Xt n|2⌘ analytical ex-pressions. LetHe and Ter¨asvirta [13] proved that, for   2 {1,2} and k a positive integer, # ,k < 1 is a necessary and su cient condition for the existence of E ⇣|Xt| k⌘.Ling and McAleer [14] note that He and Ter¨asvirta’s proofs actually hold for any   > 0 and that # ,k < 1 also guarantees the strict stationarity of X. However, only for   = 1 both autocorrelations of interest were obtained in [13].# ,k =E⇢h↵1 Zt+   + 1 Zt    + 1ik . (2.2)Thus, in the following, let us assume that X is a 1-TGARCH(1,1) (also writ- ten TGARCH(1,1)) process such that the generating process margins have zero mean, unit variance and a symmetric distribution with finite fourth moment.To simplify our notation, let #k = #1,k and  k = E ⇣|Zt|k⌘.If #2 < 1, X is strictly and weakly stationary and ⇢n (1) exists for all n, with⇢1(1)=#1 +  1 1 #21   21  1  (2.3) (1 #21)  21 (1 #2)]]></page><page Index="51" isMAC="true"><![CDATA[ -TGARCH MODEL: CAPTURING THE TAYLOR EFFECT 41 and, for n > 1,⇢ (1)=#n 1⇢ (1). (2.4) n11The su cient condition for the existence of the squared process autocorre- lations is #4 < 1, which implies #2 < 1. Under this condition,↵ 12 +   12   3   12 ⌥ ⇢1(2)= 2 +(↵1+ 1) 1  +  +(1 #4) (2.5)(2.6)and, for n > 1,where⌥=[(↵1 + 1) 3 +2 1](1+2#1 +2#2 +#1#2)(1 #1)(1 #2)n 1⇢ (2)=#n 1⇢ (2)+(1 # ) ⇤ X#n i 1#i,n214 21 i=144+ ⇣ ↵ 21 +   12 + ( ↵ 1 +   1 )   1   3 +   12   1 ⌘ ( 1 + # 1 ) 2 ( 1   # 3 ) +   1   # 21   ( 1   # 2 ) ( 1   # 3 ) , 2  4  4 ⇤ =  2#1 (1+#1) (1 #3) + [(↵1+ 1)  3 + 2 1] (1+2#1+2#2+#1#2) (1 #1) ,  =  4 (1   #1) (1   #2)   (1 + #1)2 (1   #3) (1   #4) , and=1+3#1 +5#2 +3#3 +3#1#2 +5#1#3 +3#2#3 +#1#2#3.We note that the parameter ↵0 has no interference in the existence conditions nor in the autocorrelations expressions.The expressions of the autocorrelations displayed here have been suitably adapted from [13]. These adaptations were crucial in obtaining the results pre- sented in the following section.3. The Taylor property in the 1-TGARCH(1,1) modelIn this section, we begin by establishing the presence of the Taylor property in the 1-TGARCH(1,1) model. Then, we explore the extent of the parameteriza- tions region that verifies it. Finally, considering the first-order autocorrelations and particular distributions for the generating process margins, we illustrate the influence of its kurtosis in the size of the referred region.To develop this study we assume an additional condition on the moments of the generating process, more specifically, that   1/4 <   < 1. This fundamen-tal relation, first introduced by Gon¸calves, Leite and Mendes-Lopes [5], was further analysed and then characterized by Haas [12], as “will be satisfied for pratically any distribution that one can anticipate in the context of GARCH models”.41]]></page><page Index="52" isMAC="true"><![CDATA[42 J. LEITETheorem 3.1. The 1-TGARCH(1,1) model, with symmetric, centered and reduced generating process marginal distribution, such that   1/4 <   < 1,41has a parameterization set, contained in the set defined by #4 < 1, for whichthe Taylor property is verified.Proof. Let us admit that ↵ is arbitrarily fixed and let us define the following n   0 o3   12 ( ↵ 1 +   1 )2  1,setC= (↵, , )2 R+ 3:# <1 . 11104We have#1 = ↵1 + 1 1 + 1,2↵ 12 +   12 2#2 = 2 +  1 +  1 (↵1 +  1)  1,↵ 13 +   13 3 3   1   ↵ 12 +   12   #3= 2  3+ 1+ 2 +and#4 = 2  4 + 1 +2 1 ↵1 + 1  3 +3 1 ↵1 + 1 +2 1 (↵1 + 1) 1.↵14+ 14 4  3 3  2 2 2  3We start by observing that both ⇢n (1) and ⇢n (2), for each n 2 N, as func- tions of ↵1,  1 and  1, are continuous for all points of C.Let us now consider that (↵ ,   ,   ) converges to (  1/4 ,   1/4 , 0) through 11144values in C. Then, from expression (2.3), we can guarantee that ⇢1(1)  !     1/4, since the denominator (1   #2)    2(1   # ) converges to 1    2 6= 0.Analogously, taking into account expression (2.5), we have ⇢ (2)  !   1/2, 14141121because (1   #4) converges to zero and, as the expression (1   #1)(1   #2) converges to (1     1/4)(1   1/2) 6= 0, the denominator   does not convergeto zero. Therefore, recalling that   >   1/4, we can state that     1/4 > 1414  1/2 and conclude that there exists a neighbourhood of (  1/4,   1/4, 0) that, 444intersected with C, only has points corresponding to parameterizations of the 1-TGARCH(1,1) that verify ⇢1(1) > ⇢1(2).144In a similar manner, for each n in N\ {1}, when (↵1,  1,  1) converges to (  1/4,   1/4, 0) through values in C,44 ⇣ 1⌘n ⇢ (1) = #n 1⇢ (1)  !     4 ,n1114⇤X ⇣⌘n  n 1  1⇢ (2)=#n 1⇢ (2)+(1 # ) #n i 1#i  !   2 ,n214 214 i=1and so we can conclude that there exists a neighbourhood of (  1/4,  1/4,0) 44such that, intersected with C, its elements satisfy ⇢n (1) > ⇢n (2). ⇤]]></page><page Index="53" isMAC="true"><![CDATA[ -TGARCH MODEL: CAPTURING THE TAYLOR EFFECT 43It follows easily from the proof arguments above that the parameterizationsof the 1-TGARCH(1,1) model exhibiting the Taylor property are, at least, theones with ↵ '   .   1/4 and   close or equal to zero, i.e., parameterizations 114 1close to the AVARCH(1) model and near the boundary of existence of the autocorrelations. The focus in these particular parameterizations derives from prior works, namely He and Ter¨asvirta [13] and Gonc¸alves, Leite and Mendes- Lopes [5]. We will now show that this parameterizations set can be enlarged.In order to do that, we set  1 = k↵1, where k is a real positive con-stant. Considering the same hypotheses of Theorem 3.1 and following the samereasoning of its proof, we first observe that, when (↵ ,  ,  ) converges to ✓  1/4⇣ 2 ⌘1/4  1/4⇣ 2 ⌘1/4 ◆ 1 1 1 4 k4+1 ,k 4 k4+1 ,0 through values in C, then #4 ! 1. Moreover, for each n in N, we also haveandsince the denominators of the autocorrelations converge to nonzero numbers. Thus, taking into account these limits, the Taylor property holds for values of⇢n(1) !"1+k·  1 ✓ 2 ◆1/4#n 2  1/4 k4 + 14"1 + k2  1/2 ✓ 2 ◆1/2#n⇢n (2)  ! 2 ·  4 k4 + 1 ,k > 0 such that1+k✓ 2 ◆ 1/4 11 + k2 k4 + 1 >    1/4 . (3.1)14As   1/4 <   is equivalent to (   1/4) 1 < 1, then the solutions k > 0 of 4114(3.1) are also solutions of the following inequality 1+k ✓ 2 ◆ 1/41 + k2 k4 + 1   1, (k 1)2  k6 +2k5 +3k4 +8k3 +3k2 +2k 1  0.(3.2)(3.3)which is equivalent toNumerical analysis methods allow us to determine that if k 2 [0.2814, 3.5546] then it satisfies (3.3), thus also satisfies (3.1).In view of the previous calculations, we can conclude that the Taylor prop- erty holds, at least, for parameterizations in a neighbourhood of the boundary of existence of the autocorrelations such that 0.2814↵1 <  1 < 3.5546↵1 (so ↵1 and  1 can be clearly di↵erent) and with  1 close or equal to zero, covering]]></page><page Index="54" isMAC="true"><![CDATA[44 J. LEITEwell over half the boundary when  1 = 0. Hence, the Taylor property is present in models that can also reproduce the leverage e↵ects.The extension to models with  1 not so close to zero is due to the continuity of ⇢n (1) and ⇢n (2), when viewed as functions of ↵1,  1 and  1, and arises when the results presented here are considered with the one that establishes the presence of the Taylor property in the AVGARCH(1,1) model, which was obtained by Haas [12, corollary 3].To further explore the parameterizations region satisfying the Taylor prop- erty, we now consider that the generating process variables has a centered and reduced distribution based on Student’s t-distribution with n degrees of free- dom, i.e., with density1   n+1  ✓ x2 ◆ (n+1)/2fZ (x) = p   2   1 + , x 2 R.Considering 120, 30, 15, 14, 7 and 5 degrees of freedom, so moving away from the Gaussian distribution, and using expressions (2.3) and (2.5) with  1 = 0, in Figure 1 we have identified parameterizations regions that satisfy the relation ⇢1 (1) > ⇢1 (2) for the 1-TARCH(1) model.In Figure 1 we can observe that the region where ⇢1 (1) > ⇢1 (2) is verified seems to enlarge as the kurtosis, kZt , of the generating process variables in- creases, since we have, for the decreasing values of the considered degrees of free- dom, kZt equal to 3.05, 3.23, 3.55, 3.6, 5 and 9, respectively. It is also relevant to notice that when kZt is equal 5 and 9, almost all, if not all, parameterizations for which the existence of the autocorrelations is ensured, verify ⇢1 (1) > ⇢1 (2). These representations were obtained for other centered and reduced distribu- tions, namely, triangular (kZt =2.4), Gaussian (kZt =3) and Laplace (kZt =6), displaying the same behaviour.4. The Taylor effect in the  -TGARCH(1,1) modelIn this section, we consider the first-order  -TGARCH model with   not necessarily equal to 1. Since both theoretical expressions of ⇢1 (1) and ⇢1 (2) are only available for   = 1, the Taylor e↵ect is analysed here via simulation.Taking into account the final remarks of the previous section, we considered two zero mean and unit variance distributions for the generating process mar- gins, namely the Gaussian and Laplace. For these distributions and in order to compare the ability that models with di↵erent values of   have to gener- ate series with the Taylor e↵ect, we considered four values of  , specifically 0.25, 0.5, 1 and 2, and three parameterizations for each one of them. In order to guarantee the existence of the autocorrelations, all parameterizations satisfy # ,4/  < 1, with 4/  an integer. These eight sets of three parameterizations have(n 2)⇡   n n 2 2]]></page><page Index="55" isMAC="true"><![CDATA[Figure⇢1 (2)verification region for the1- -TGARCH MODEL: CAPTURING THE TAYLOR EFFECT45Student df 120 1.00.8 0.6 0.4 0.2 0.00.00.20.40.60.81.0Α1Student df 15 1.00.8 0.6 0.4 0.2 0.00.00.20.40.60.81.0Α1Student df 7 1.00.8 0.6 0.4 0.2 0.00.00.20.40.60.81.0 Α1Student df 30 1.00.8 0.6 0.4 0.2 0.00.00.20.40.60.81.0Α1Student df 14 1.00.8 0.6 0.4 0.2 0.00.00.20.40.60.81.0Α1Student df 5 1.00.8 0.6 0.4 0.2 0.00.00.20.40.60.81.0 Α1Β1 Β1 Β1Β1 Β1 Β11. ⇢1 (1) >TARCH(1) model with generating process marginal distribu- tion based on Student’s t-distribution with df degrees of free- dom.approximately the same values of the model theoretical kurtosis. Our choice of comparing the models with the same kurtosis is related to the findings of He and Ter¨asvirta [13] for the AVGARCH(1,1) model, who calculated the theoreti- cal kurtosis expression and observed that its high values favours the appearance of the Taylor e↵ect. For each parameterization, 100 trajectories with 4000 ob- servations were obtained, with the first 1000 observations discarded to remove the e↵ect of the initial values choice. For each simulated trajectory, the sam- ple autocorrelation of the absolute observations was compared with the one of squared observations. The simulation results are presented in Table 1 for the Gaussian distribution and in Table 2 for the Laplace distribution.]]></page><page Index="56" isMAC="true"><![CDATA[46J. LEITE  ↵1  1 K0.25 0.51 0.3 0.8 38.55 810.5 0.7 0.61 0.5 0.53 0.530.3 0.75 0.65 0.3737.33 98 11.46 97 6.97 9839.49 97⇢ˆ1 (1) > ⇢ˆ1 (2) Score (out of 100)11.73 91 0.545 0.545 6.94 9320.7 0.1 0.53 0.530.3 0.686 0.68 0.1 0.44 0.4411.60 58 6.57 6638.51 28 11.21 11 6.62 9Table 1. Simulation results considering the  -TGARCH(1,1) model with ↵0 = 1,  1 = 0.1, Zt ⇠ Gaussian and  , ↵1,  1 specified above and K is the model theoretical kurtosis.  ↵1  1 K⇢ˆ1 (1) > ⇢ˆ1 (2) Score (out of 100)0.25 0.46 0.350.2 0.5 0.20.4 0.21 0.2 0.470.22 0.10.35 0.15Table 2. Simulation results considering the  -TGARCH(1,1) model with ↵0 = 1,  1 = 0.1, Zt ⇠ Laplace and  , ↵1,  1 specified above and K is the model theoretical kurtosis.0.6 0.2 0.20.62 0.05 0.20.645 0.1 0.238.73 99 11.44 99 6.91 9838.82 93 11.94 83 6.73 9537.59 82 11.2 78 6.68 790.50.190.15 6.82 3533.6 45 11.25 31]]></page><page Index="57" isMAC="true"><![CDATA[ -TGARCH MODEL: CAPTURING THE TAYLOR EFFECT 47Analysing Tables 1 and 2 separately, it appears that increasing values   are not favourable to Taylor e↵ect, especially for   = 2, which includes the standard GARCH model. Looking at the model kurtosis, it seems that it plays a role in the appearance of the Taylor e↵ect, larger values being preferable, but other factors also seem to be at play. Now comparing both tables regarding the generating process marginal distribution, the leptokurtic Laplace distribution apparently generates more series with this e↵ect than the mesokurtic Gaussian distribution. Although these results do not contradict previous findings, further studies are required to fully support these claims.AcknowledgmentsI would like to thank the referee for helpful comments and suggestions.I kindly acknowledge the precious comments and revision of Professor Nazar´e Mendes Lopes and Professor Esmeralda Gon¸calves in the development of this theme, as my PhD advisors until 2013. In a broader sense, I am grateful to them for introducing me to time series analysis and for their supreme guidance in so many stages of my academic training. On this occasion, I especially wish to thank Professor Nazar´e Mendes Lopes for sharing her knowledge, insights and uplifting way to approach problems. I have to say... it has been a privilege.References[1] F. Black, Studies of stock price volatility changes, in: Proc. of the 1976 Meeting of the Business and Economic Statistics Section, American Statistical Association, 177–181, 1976.[2] Z. Ding, C. W. J. Granger, and R. F. Engle, A long memory property of stock market returns and a new model, J. Empir. Financ. 1, 83–106, 1993.[3] R. F. Engle, New frontiers for ARCH models, J. Appl. Econometrics 17, 425–446, 2002.[4] C. Francq and J. M. Zakoian, GARCH Models: Structure, Statistical Inference andFinancial Applications. John Wiley & Sons, Chichester, UK, 2010.[5] E. Gonc¸alves, J. Leite, and N. Mendes-Lopes, A mathematical approach to detect theTaylor property in TARCH processes, Statist. Probab. Lett. 79, 602–610, 2009.[6] E. Gon¸calves, J. Leite, and N. Mendes-Lopes, On the probabilistic structure of power threshold generalized ARCH stochastic processes, Statist. Probab. Lett. 82, 1597–1609,2012.[7] E. Gon¸calves, C. Martins, and N. Mendes-Lopes, The Taylor property in non-negativebilinear models, REVSTAT (in press), 2015.[8] E. Gon¸calves, N. Mendes-Lopes, I. Dorotoviˇc, J. M. Fernandes, and A. Garcia, Northand South Hemispheric Solar Activity for Cycles 21-23: Asymmetry and ConditionalVolatility of Plage Regions Areas, Sol. Phys. 289 (6), 2283–2296, 2014.[9] C. W. J. Granger, The past and future of empirical finance: some personal comments,J. of Econometrics 129, 35–40, 2005.[10] C.W.J.GrangerandZ.Ding,Somepropertiesofabsolutereturn:analternativemeasureof risk, Annales d’E´conomie et de Statistique 40, 67–91, 1995.]]></page><page Index="58" isMAC="true"><![CDATA[48 J. LEITE[11] C. W. J. Granger, S. Spear, and Z. Ding, Stylized facts on the temporal distributional properties of absolute returns: an update, in: Proceedings of the Hong Kong International Workshop on Statistics and Finance: An Interface, W. Chan, W. Li, H. Tong (eds.), Imperial College Press, London, 97–120, 2000.[12] M. Haas, Persistence in volatility, conditional kurtosis, and the Taylor property in ab- solute value GARCH processes, Statist. Probab. Lett. 79 (5), 1674–1683, 2009.[13] C. He and T. Ter¨asvirta, Properties of moments of a family of GARCH processes, J. Econometrics 92, 173–192, 1999.[14] S. Ling and M. McAleer, Stationarity and the existence of moments of a family of GARCH processes, J. Econometrics 106, 109–117, 2002.[15] H. Malmsten and T. Tera¨svirta, Stylized facts of financial time series and three popular models of volatility, Eur. J. Pure Appl. Math. 3, 443-477, 2010.[16] B. Mandelbrot, The variation of certain speculative prices, J. Bus. 36, 394–419, 1963.[17] A. Mora-Gala´n, A. P´erez, and E. Ruiz, Stochastic volatility models and the Taylor e↵ect, Statistics and Econometrics Series Working Papers ws046315, Universidad Carlos III deMadrid, 2004.[18] A. P´erez, E. Ruiz, and H. Veiga, A note on the properties of power-transformed returnsin long-memory stochastic volatility models with leverage e↵ect, Comput. Statist. DataAnal. 53, 3593–3600, 2010.[19] S. J. Taylor, Modelling Financial Time Series, Jonh Wiley & Sons, Chichester, UK,1986.[20] S. J. Taylor, Asset Price Dynamics, Volatility, and Prediction. Princeton UniversityPress, Princeton, 2005.[21] H. Veiga, Financial Stylized Facts and the Taylor-E↵ect in Stochastic Volatility Models,Economics Bulletin 29, 265–276, 2009.(J. Leite) CMUC and Instituto Polite´cnico de Coimbra Coimbra Business School (ISCAC)Quinta Agr´ıcola - Bencanta3040-316 CoimbraPortugalE-mail address: jleite@iscac.pt]]></page><page Index="59" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION: A REVIEWM. MANUELA NEVESDedicated to Nazar´e Lopes as a token of friendshipAbstract. This paper presents an overview of some applications of re- sampling computer intensive methodologies, like the bootstrap and the jackknife, in a reliable estimation of the extremal index. The extremal index, a measure of clustering of extreme events, is a key parameter in extreme value theory in a dependent set-up. Its estimation has been con- sidered by several authors but some di culties still remain. Most of the semi-parametric estimators of this parameter show the same type of be- haviour: nice asymptotic properties, but a high variance for small k, the number of upper order statistics used in the estimation; a high bias for large k and the need for an adequate choice of k. After a brief review of some estimators and their asymptotic properties, two computational procedures, the Block-Bootstrap and Jackknife-After-Bootstrap are ap- plied to improve the extremal index estimation. An adaptive choice of the block size for the block-bootstrap resampling is presented. A few results of a simulation study will illustrate the application of that choice.1. Introduction and the scope of the paperIn Extreme Value Theory (EVT) we deal essentially with the estimation of parameters of extreme or rare events. A large number of applications in areas such as biology, environment, finance, hydrology and telecommunications, reveals the importance of adequate estimation procedures.The classical assumption in EVT is that we have a set of independent and identically distributed (i.i.d.) random variables (r.v.’s), X1, . . . , Xn, from anAccepted: 07 April 2015.2010 Mathematics Subject Classification. Primary 62G09, 62E20, 60G10; Secondary60G70.Key words and phrases. Bootstrap and Jackknife, extremal index, semi-parametric esti-mation, statistics of extremes.The work was supported by National Funds through FCT – Fundac¸˜ao para a Ciˆencia e aTecnologia, project PEst-OE/MAT/UI0006/2014 (CEAUL).49]]></page><page Index="60" isMAC="true"><![CDATA[50 M. M. NEVESunknown distribution function (d.f.) F and we are concerned with the lim- iting behaviour of either Mn ⌘ Xn:n = max(X1,...,Xn) or mn ⌘ X1:n = min(X1, . . . Xn) as n ! 1. Dealing with maximum values, whenever it is pos- sible to linearly normalize Mn so that we get a non-degenerate limit, such a limit is the Extreme Value (EV) d.f., given byEV⇠(x):=⇢ exp[ (1+⇠x) 1/⇠], 1+⇠x>0, if ⇠6=0, (1.1) exp[  exp( x)], x 2 R, if ⇠ = 0.We then say that F is in the domain of attraction for maxima of EV⇠, denoting this by F 2 DM (EV⇠ ).The EV⇠ d.f., in (1.1), incorporates the three Fisher-Tippett types: the Gum- bel family, ⇤(x) ⌘ EV0(x) := exp( exp( x)), x 2 R, (⇠ = 0), the limit for exponentially-tailed distributions; the Fr´echet family,  ↵(x) ⌘ EV1/↵(↵(x  1)) := exp( x ↵), ↵ > 0, x > 0, (⇠ = 1/↵ > 0), the limit for negative polynomial heavy-tailed distributions and the Weibull family, ↵(x) ⌘ EV 1/↵ (↵(x + 1)) := exp( ( x)↵), ↵ > 0, x < 0, (⇠ =  1/↵ < 0), the limit for short-tailed distributions. If ⇠ = 0, the right endpoint, x⇤ := sup{x : F(x) < 1}, can then be either finite or infinite. If ⇠ > 0, F has an infinite right endpoint. If ⇠ < 0, F has a finite right endpoint. The shape parameter, ⇠, is directly related to the weight of the right tail, F := 1   F , of the underlying model F . As ⇠ increases the right tail becomes heavier.Whenever independence is no longer valid, some important dependent se- quences have been studied. The limit distributions of their order statistics, under some dependence structures, have been obtained. Stationary sequences are examples of such sequences and are realistic for many real problems.As dependence in stationary sequences can assume several forms, some con- ditions have to be imposed. The first condition, known as the D(un) dependence condition, Leadbetter et al. (1983), ensures that any two extreme events can become approximately independent as n increases when separated by a rela- tively wide interval of length ln = o(n). Hence, D(un) limits the long-range dependence between such events.Under adequate local dependence conditions, the limiting d.f. of the max- imum of the stationary sequence may be directly related to the maximum of the i.i.d. associated sequence, through a new parameter, the so-called extremal index, usually denoted by ✓. This parameter also a↵ects other parameters of extreme events, so its reliable estimation is of great importance.However, estimators of ✓ show the usual behaviour of semi-parametric tail estimators: nice asymptotic properties, but a high variance for small k, the]]></page><page Index="61" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 51number of upper order statistics used in the estimation, and a high bias for large k.Resampling methodologies have provided very fruitful results in the field of statistics of extremes. The bootstrap methodology has been used, in partic- ular, in the choice of the number k of upper order statistics to be taken in the semi-parametric tail index estimation, see Draisma et al. (1999), Danielson et al. (2001), Gomes and Oliveira (2001), Gomes et al. (2012, 2013), Caeiro and Gomes (2014) and Gomes et al. (in press), to mention a few works on boot- strap in statistics of extremes. Bootstrapped version of the estimators usually shows a more stable path which helps the estimation of the optimal sample fraction, through some stability criterion, see for example Gomes and Pestana (2007). The jackknife methodology allows us to obtain the bias and the variance of a given statistic and then to build estimators with bias and Mean Square Error (MSE) smaller than those of initial estimators. Section 2 of this arti- cle is dedicated to define the extremal index and to describe some properties. Some estimators are also presented. Resampling procedures are introduced in Section 3 where attention is specially given to the dependent set-up. Classi- cal bootstrap procedures, derived for independence, have to be adapted to the dependent context. One of the procedures proposed for the bootstrap resam- pling in this situation, see for example Lahiri (2002), consists of defining blocks for resampling, instead of resampling the individual observations. But the per- formance of the bootstrap estimator crucially depends on the block size that must be supplied by the user. So, a computational procedure for estimating that block size is proposed, following Lahiri et al. (2007). A brief illustration of a simulation study that is now in progress is shown in Section 4. The paper finishes with a few comments and some notes regarding the work in progress.2. The Extremal Index: definition and estimationThe extremal index, ✓, measures the relationship between the dependence structure of the data and the behaviour of the exceedances over a high threshold un. This threshold un is such that, for ⌧ > 0, the underlying d.f. F verifiesF(un) = 1   ⌧/n + o(1/n), n ! 1. (2.1)A stationary sequence {Xn}n 1 from an underlying model F is said to have an extremal index ✓ (0 < ✓  1) if, for all ⌧ > 0, we can find a sequence of levels un = un(⌧) such that with {Yn}n 1 the associated i.i.d. sequence (i.e., an i.i.d. sequence from the same F ), it is verifiedP(Yn:n  un) = Fn(un)  ! e ⌧ and P(Xn:n  un)  ! e ✓⌧, n!1 n!1(see Leadbetter et al. (1983)).]]></page><page Index="62" isMAC="true"><![CDATA[52 M. M. NEVESProvided that a stationary sequence {Xn}n 1 has limited long-range de- pendence at extreme levels, the maxima of this sequence follow the same dis- tributional limit law as the associated independent sequence, {Yn}n 1, but with other values for the parameters. This result was established by Leadbetter (1983, Theorem 2.5), and fully detailed in the proof and comments there included, see also Beirlant et al. (2004, page 377). It is presented in the following theorem.Theorem 2.1. If {Xn}n 1 is a stationary sequence with marginal distributionF, {Yn}n 1 an i.i.d. segquence of r.v.’s with the same distribution F, Mn :=max(X ,··· ,X ) and Mg:= max(Y ,··· ,Y ), under the D(u ) condition, 1nnn1onnwithun =anx+bn,Pr (Mn  bn)/an x  ! G1(x)asn !1,fornor- n!1malizing sequences {an > 0} and {bn }, if and only if, P r {(Mn   bn )/an  x}  ! G2(x) where G2(x) = G✓1(x), for a constant ✓ such that 0 < ✓  1.n!1So, given that G1(·) ⌘ EV⇠(·), the limit law G2(·) ⌘ EV⇠✓(·) is an extremevalue d.f. with location, scale and shape parameters (μ✓, ✓,⇠✓) given by 1 ✓⇠ ⇠μ✓=μ   ⇠ ,  ✓= ✓ and ⇠✓=⇠,where (μ,  , ⇠) are the location, scale and shape parameters, respectively, of the limit law of the i.i.d sequence.The extremal index ✓, 0 < ✓  1, is directly related to the clustering of ex- ceedances: ✓ = 1 for i.i.d. sequences and ✓ ! 0 whenever dependence increases. Definitions can be given extending values of ✓ to the case ✓ = 0, but this is a situation out of interest, see Beirlant et al. (2004).For illustration of the behaviour of a stationary process for some values of ✓ let us consider the following examples:Example 2.2. Let {Xn}n 1 be the two-dependent sequence defined as Xn = max(Zn+1,Zn), n   1, where {Zi}i 1 are standard exponential i.i.d. r.v.’s. The underlying model for Xn is given by F (x) = (1   exp( x))2, x   0.Let {Yn}n 1 be a sequence of i.i.d. r.v.’s from the same distribution, i.e., F(y) = (1   exp( y))2, y   0.Figure 1 shows a size equal to 2 for the clusters of exceedances of high levels by the {Xn} sequence. Actually for that sequence we have ✓ = 1/2. It can also be seen a shrinkage of the largest observations for the 2-dependent sequence, despite of the fact that we have the same model underlying both sequences.Example 2.3 (Beirlant et al., 2004, Example 10.3). “Max-Autoregressive Pro- cess (ARMAX process)”.]]></page><page Index="63" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 53 Let {Zi}i 1 be a sequence of independent, unit-Fr´echet distributed randomvariables. For 0 < ✓  1, letX1 =Z1 Xi =max{(1 ✓)Xi 1,✓Zi}, i 2.Forun =nx,0<x<1,Pr Mn un !exp  ✓/x ,asn!1,sothe extremal index of the sequence is ✓.iid 2dep0 10 20 30 40 kFigure 1. One realization of an i.i.d. process and a 2-dependent process.17 1.072552 0.686968 3.097931 18 0.349065 1.948675 2.78813819 0.496704 0.974338 2.509324 20 20 1.823144 0.779399 2.258392 15 21 0.817736 0.964604 6.4018122 0.68322 1.447588 5.761629 10 23 0.453179 1.237074 5.185466 5 24 5.197649 0.618537 4.66692 25 4.337472 0.369991 4.200228=0.90 26 0.614502 0.560788 3.7802050 10 20 30 40 5027 0.692333 0.684494 3.40218420 15 105=0.500 10 20 30 40 50=0.1 2015 10 5 00 10 20 30 40 50Figure 2. One realization of an ARMAX process with ✓ = 0.9; 0.5 and 0.1.Figure 2 shows a partial realization of variables following an ARMAX process with ✓ = 0.9; 0.5 and 0.1, respectively. The maxima show increasing clustering as ✓ ! 0. Notice also again a ‘shrinkage of maximum values’ as dependence increases.One of the local dependence conditions guaranteeing the existence of an extremal index is the D00 condition, introduced by Leadbetter and Nandagopalan (1989), under which it can be given a very common interpre- tation of ✓, as being the reciprocal of the ‘mean time of duration of extreme012345]]></page><page Index="64" isMAC="true"><![CDATA[54 M. M. NEVESevents’ which is directly related to the exceedances of high levels (Hsing et al., 1988, Leadbetter and Nandagopalan, 1989),✓=1. limiting mean size of clustersIdentifying clusters by the occurrence of downcrossings or upcrossings, we can write✓= limPr[X2 un|X1 >un]= limPr[X1 un|X2 >un]. (2.2) n!1 n!1Given a sample, (X1, . . . , Xn), the empirical counterpart of the above inter- pretation led to the naive estimators: the classical up-crossing (down-crossing)estimator, UC estimator , ⇥bUC (DC-estimator, ⇥bDC) (Nandagopalan, 1990; Gomes, 1990, 1992, 1993)n 1P I(Xi un <Xi+1)⇥bUC(un) := i=1I(Xi > un)for a suitable threshold un, where I(A) denotes, as usual, the indicator function of A. Consistency of this estimator is obtained provided that the high level un is a normalized level, i.e. if with ⌧ ⌘ ⌧n > 0, the underlying d.f. F verifiesF(un)=1 ⌧/n+o(1/n), n!1 and ⌧/n!0.Di↵erent forms of identifying clusters gave rise to other estimators. Let us men- tion two very popular estimators: the blocks estimator and the runs estimator, Hsing (1991, 1993). The blocks estimator is derived by dividing the data set into approximately kn blocks of length rn, where n ⇡ kn ⇥ rn, i.e., considering kn = [n/rn]. Each block is treated as one cluster and the number of blocks in which there is at least one exceedance of the threshold un is counted. The blocks estimator, ⇥bBn (un), is then defined asPkn I max X ,···,X  >u   ⇥bB(u):= i=1 P(i 1)rn+1 irn n.n n ni=1 I (Xi > un)If we assume that a cluster consists of a run of observations between twoexceedances, then the runs estimator is defined as:Pn i=1n 1P I(Xi >un,Xi+1 un)⌘ i=1 := ⇥bDC(un)I(Xi > un)(2.3)⇥n(un):= Pni=1I(Xi >un) .Pn i=1bR Pni=1 I  Xi > un,max Xi+1,··· ,Xi+rn 1   un Under mild conditions, limn!1 ⇥bBn (un) = limn!1 ⇥bRn (un) = ✓. Other proper- ties of these estimators have been well studied by Smith and Weissman (1994) and Weissman and Novak (1998).]]></page><page Index="65" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 55In this paper our attention will be focused on the UC-estimator in (2.3). Given the sample (X1,...,Xn) and the associated ascending order statistics, X1:n  · · ·  Xn:n , we shall consider the deterministic level u ⌘ un substituted by the stochastic one, Xn k:n, and write the UC estimator, in (2.3), as a function of k (Gomes et al., 2008),b U C b U C 1 nX  1⇥ ⌘ ⇥ (k):= k I(Xi Xn k:n <Xi+1). (2.4)i=1For many dependent structures, the bias of ⇥bUC(k) has two dominant com-(2.5)ponents of orders k/n and 1/k (see Gomes et al., 2008), Bias[⇥bUC(k)]='1(✓)k +'2(✓)1 +o✓k◆+o✓1◆,nknk whenever n ! 1 and k ⌘ k(n) ! 1, k = o(n).Bias22 kkkFigure 4. MSE, Var and Bias2 of ⇥bUC(k) for ✓ = 0.9,0.5,0.1 for the ARMAX process in Example 2.Figure 3. Mean values of ⇥bUC(k) for ✓ = 0.9,0.5,0.1 for the ARMAX process in Example 2.MSEVar VarBiasMSEBias2MSEVar]]></page><page Index="66" isMAC="true"><![CDATA[56 M. M. NEVESFigures 3 and 4 illustrate the behaviour of the bias, variance and MSE of the ⇥bUC(k) estimator.Remark 2.4. Some remarks on Figures 3 and 4:• The ⇥bUC estimator shows a very strong bias.• The bias is the dominant component of the MSE.• MSE(⇥bUC) is very sharp, which reveals a need for a very accurate way of choosing k in order to obtain a reliable estimate of ✓. Alternatively it would be sensible to look for less biased estimators.Recently resampling techniques have been used, specifically Bootstrap and Jackknife procedures, which have shown themselves to improve the performance of the semi-parametric estimators. Let us refer to some recent works deal- ing with those procedures to estimate parameters in EVT, such as Draisma et al. (1999), Gomes and Oliveira (2001), Danielsson et al. (2001), Gomes et al., (2011, 2012), to cite only a few. Prata Gomes and Neves (2011, 2015a,b) present some results on the use of resampling procedures in the extremal index estimation.3. Resampling Techniques and the dependent set-upResampling methodologies, Efron (1979), have revealed good results in the estimation of the threshold k, and in the reduction of bias of any estimator of a parameter of extreme events. In its classical form, the bootstrap has proven to be a powerful nonparametric tool when based on i.i.d observations. But Singh (1981) showed that it could be inadequate under dependence.So the bootstrap methodology needs to take into account the two di↵erent situations: resampling from an i.i.d. sequence or resampling from a dependent sequence.3.1. The Bootstrap under dependence. Several attempts have been made to extend the bootstrap method to the dependent case. A breakthrough was achieved when resampling of single observations was replaced by block resam- pling.Hall(1985), Carlstein (1986), Ku¨nsch (1989) and Liu and Singh (1992) have independently introduced nonparametric versions of the Bootstrap and Jack- knife applicable to weakly dependent stationary observations. Their resampling technique considers to resample or to delete one-by-one whole blocks of ob- servations to obtain consistent procedures for estimating a parameter of the distribution of the stationary series.]]></page><page Index="67" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 57The motivation for this scheme is to preserve the dependence structure of the underlying model within each block. Several ways of blocking have been proposed: Nonoverlapping Block Bootstrap (NBB), Carlstein (1986); Moving Block Bootstrap (MBB), Ku¨nsch, (1989), Liu and Singh, (1992); Circular Block Bootstrap (CBB), Politis and Romano (1992) and Stationary Bootstrap (SB), Politis and Romano (1994).For each way of blocking it is necessary to consider a length b ⌘ b(n) to resample blocks of observations, but the accuracy of block bootstrap estimators depends critically on the block size for resampling, that must be supplied by the user, (Lahiri et. al., 2007).Figure 5 illustrates one sample path of size n = 100 generated from the ARMAX process with ✓ = 0.1 and also sample paths obtained from one classical bootstrap resampling and block bootstrap resampling using several block sizes. It is clear that the block size a↵ects strongly the pattern of the resampled values.9  7 8 7 6 5 4  3  3 2 1 6  5  4 2 1 0  0 0  20  40  60  80  100  120  0  20  40  60  80  100  120 9  9 8  8 7  7 6  6 5  5 4  4 3  3 2  2 1  1 0  0 0  20  40  60  80  100  120  0  10  20  30  40  50  60  70  80  90  100 Figure 5. Values from a sample of size n = 100 generated from the ARMAX process with ✓ = 0.1 (top left) and resampled equal size samples considering the classical i.i.d. bootstrap (top right) and blocks of size 5 and 15 (bottom left and right, respectively).Several authors such as Hall et al. (1995), Bu¨hlman and Ku¨nsch (1999), Politis and White (2004) and Lahiri et al. (2007) proposed ways of estimating the optimal block size.Here we follow Lahiri et al. (2007), who proposed a nonparametric plug-in (NPPI) method for the empirical choice of the optimal block size for the block]]></page><page Index="68" isMAC="true"><![CDATA[58 M. M. NEVESbootstrap estimation of characteristics of an estimator. Two of these charac- teristics are the bias and the variance of the estimator. This method employs non-parametric resampling procedures to estimate the relevant constants in the leading term of the “optimal” block length, so it does not require the knowledge and/or derivation of explicit analytical expressions for those constants.Extremal index estimators usually have a high bias. In most cases the bias is the main component of the MSE. As a first step it was thought to control the bias through a bootstrap estimator, that needs an adequate block size choice for resampling. The method applied here is based on the Jackknife-After-Bootstrap (JAB) of Efron (1992) and Lahiri (2002) to estimate the relevant constants that appear in expansions of the bias of the bootstrap estimator, in order to obtain a more stable and accurate estimate path.3.2. The NPPI procedure. Given the observations (X1,X2,...,Xn), let ⇥bn be an estimator of ✓. Let us consider now the bias of ⇥bn n ⌘Bias⇣⇥bn⌘=E⇣⇥bn⌘ ✓,and the corresponding block bootstrap estimator based on blocks of size b (weshall consider here the MBB), b ⇤ ( b ) = E ⇣ ⇥b ⇤ ( b ) ⌘   ✓b ,where E⇤ denotes the conditional mean, given the data.Lahiri et al. (2007), Section 2, remember that for many population param-eters (denoted here by  n), the variance of the corresponding block bootstrap estimator is an increasing function of the block size, b, while its bias is a decreas- ing function of b, and that under suitable regularity conditions, the variance and the bias of the block bootstrap estimator admit expansions of the formn2a Var⇣ b⇤n(b)⌘ = C1n 1b+o n 1b , (3.1)na Bias ⇣ ⇤n(b)⌘ = C2b 1 + o  b 1  , (3.2)n⇤nnas n ! 1 and over a suitable set n ⇢ {2,··· ,n} of possible block length b. C1, C2 and a are constants depending on the characteristics under study. Hall et al. (1995) and Lahiri et al. (2007) showed that the optimal block size b0n has the formb0n = ⇣2C2 ⌘1/3n1/3(1 + o(1)). C1Now, for obtainingC1 ⇠nb 1Var⇣ b⇤n(b)⌘ and C2 ⇠bBias⇣ b⇤n(b)⌘,]]></page><page Index="69" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 59they propose a consistent estimation of Var⇣ b⇤n(b)⌘ and Bias⇣ b⇤n(b)⌘ given by d[b⇤ d⇣b⇤b⇤⌘Varn ⌘ VARJAB( n(b)), Biasn = 2  n(b)    n(2b) , [where VARJAB is defined in (3.3). Parameters C1 and C2 can be consistently estimated byb ⇣b⇤ b⇤ ⌘ C1 = nb VARJAB( n(b)), C2 = 2b  n(b)    n(2b) .b  1[ b⇤The NPPI estimator of the optimal block size is then given byb0n = 2Cb2 1/3n1/3. ⇣ b2⌘C1The JAB methodology, allowing to assess the accuracy of bootstrap estima- tors for dependent data, was derived by Lahiri (2002). The key step is to delete resampled blocks instead of blocks of original data values.If  b⇤n(b) is the MBB estimator of  n and ` = n b+1 the number of “observable” blocks of length b:• Let m be an integer such that m ! 1 and m/n ! 0 as n ! 1, denoting the number of bootstrap blocks to be deleted.• Write M = ` m+1 and for i = 1,···,M let us define the set Ii = {1,··· ,`}\{i,··· ,i+m 1}, denoting the index set of all blocks obtained by deleting the m blocks.• Resample [n/b] from the reduced collection {Bj : j 2 Ii} and compute the ith jackknife block-deleted estimate,  b⇤i ⌘  b⇤i(b), i = 1,··· ,M.nn The JAB variance estimator for  n(b) is then defined asb⇤ XmM[ b⇤ e⇤i b⇤ 2VARJAB( n(b)) = (`   m)M ( n (b)    n(b)) , (3.3) i=1e⇤i b 1 h b⇤ b⇤i iwhere  n (b) = m ` n(b)   (`   m) n (b) is the ith block - deleted jackknifepseudo-value of  n(b), i = 1,··· ,M.Figure 6 shows sample paths for ✓bUC estimates and the corresponding blockbootstrap estimates, for an initial block size and the “‘optimal” block size. According to suggestions given in Lahiri et al. (2007) to obtain Cb1 and Cb2, as a first approach we used b = c1n1/5, m = c2n1/3b2/3 with c1 = 1 and c2 = 1.]]></page><page Index="70" isMAC="true"><![CDATA[60M. M. NEVESestb_com b=3teta estestb_com_bopt1.2 1 0.8 0.6 0.4 0.2 01.2 1 0.8 0.6 0.4 0.2 01.2 1 0.8 0.6 0.4 0.2 0Figure 6. One sample path of the estimator ⇥bUC for a sample of size n = 500 from an ARMAX process. Block bootstrap estimates using initial block size b1 = n1/5 = 3 and block bootstrap estimates using the“optimal” block length b0n (✓ = 0.1; bopt = 111) (✓ = 0.5; bopt = 43) and (✓ = 0.9; bopt = 38).0 10 20 30 40teta estestb_com b=3estb_com_bopt0 50 100 150 200teta estestb_com b=3estb_com_bopt0 50 100 150 200 250 300]]></page><page Index="71" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 613.3. The Generalized Jackknife methodology. The Generalized Jackknife methodology, Gray and Schucany (1972), has the properties of estimating the bias and the variance of any estimator, leading to the development of estimators with bias and mean squared error often smaller than those of an initial set of estimators.Using the information on the bias of the extremal index estimator ⇥bUC, in (2.4), Gomes et al. (2008) considered first a generalized jackknife estimator of order 2, based on ⇥bUC computed at the three levels, k, bk/2c+1 and bk/4c+1, where bxc denotes, as usual, the integer part of x. They got the estimator⇥bGJ ⌘ ⇥bGJ(k):=5⇥bUC([k/2]+1) 2 ⇥bUC([k/4]+1)+⇥bUC(k) . (3.4)This is an asymptotically unbiased estimator of ✓, in the sense that it can remove the two dominant components of bias referred to in (2.5).More generally, Gomes et al. (2008) considered the levels k, b kc + 1 and b 2kc + 1, depending on a tuning parameter  , 0 <   < 1, and the class of estimators,bGJ( ) ( 2 + 1) ⇥bUC ([ k] + 1)     ⇣⇥bUC  [ 2k] + 1  + ⇥bUC(k)⌘⇥ (k) := (1    )2 . (3.5)Actually ⇥bGJ(1/2)(k) ⌘ ⇥bGJ, given in (3.4). Among the members of the class in (3.5), those authors have been heuristically led to the choice   = 1/4. Distri- butional properties of ⇥bGJ(1/4)(k) have been obtained by Gomes et al. (2008) through simulation techniques, see also Gomes et al. (2015).In this paper, the estimator ⇥bGJ in (3.4) will be considered for illustrating the simulation study.Remark 3.1. From several studies, Gomes et al. (2008), Neves et al. (2015), Prata Gomes and Neves (2015a, 2015b), some remarks can be pointed out regarding these estimators:• The Generalized-Jackknife estimator, ⇥bGJ, shows a more stable sim- ulated mean value, near the target value of the parameter but at ex- penses of a very high variance. This does not enable it to outperform the original estimator, regarding MSE at optimal levels.• MSE(⇥bGJ) is not so sharp as MSE(⇥bUC), suggesting less dependence on the value k for obtaining the estimate of ✓.]]></page><page Index="72" isMAC="true"><![CDATA[62 M. M. NEVES4. A few simulation resultsAn extensive simulation study is being carried out by the author and some results will appear in Prata Gomes and Neves (2015a, 2015b). Here some re- marks and overall comments are given. A brief description of what is being done follows.Simulated samples are generated from a given process, with ✓ known. A simple path of the estimators ⇥bUC and ⇥bGJ is generated. Through the proce- dures briefly described in Section 3.2 the optimal block length is obtained for each sample and block bootstrap estimates, using that optimal block length, are calculated.1.8  UC  1.6  GJ  1.4 1.2 1  0.8  0.6  0.4  0.2  0 0  100 UC_Bopt=111  GJ_Bopt=60 1.8  UC  1.6  GJ  1.4 1.2 1  0.8  0.6  0.4  0.2  0 UC_Bopt=100  GJ_Bopt=47 200 300  400 1.4 1.2  GJ 1  0.8  0.6  0.4  0.2  0 500 0  100  200 UC_Bopt=6  GJ_Bopt=7 300  400  500 UC 0  100 200 300  400  500 Figure 7. One sample path of the estimators ⇥b U C and ⇥b GJ for a sample of size n = 500 from the ARMAX process and block boot- strap estimates using “optimal” block length for ✓ = 0.1, ✓ = 0.4 and ✓ = 0.8.]]></page><page Index="73" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 63A sketch of the simulation procedure illustrated in Figure 7 is:• A random sample of size n = 500 was generated from the ARMAXprocess with ✓ = (0.1, 0.4, 0.8);• An initial block size, b = n1/5 was defined, see Section 3.2.• The NPPI method was applied for estimating the “optimal” block length for resampling, which depends on the value of k.• The mode of the block size was adopted as the “optimal” block length.• Finally block bootstrap estimates for the ⇥bUC and ⇥bGJ were obtained.5. Some concluding remarksIn the field of statistics of extremes, resampling methodologies have revealed themselves as providing very good results to adequately choose the number k of upper order statistics to be taken in the semi-parametric tail index estimation or to obtain more stable path estimates, around the target value.Resampling blocks in a dependent set-up instead of resampling individual observations was proposed by some authors. However the accuracy of the es- timation strongly depends on the block size. A computational procedure for estimating the “optimal” block length for resampling in the situation of depen- dence was reviewed in this paper. In a first stage it was only considered to deal with the bias of the estimator. A simulation study for estimating the extremal index through two well known estimators, the Up-crossing and the Generalized Jackknife, was conducted.The variance of the estimators is another characteristic that needs to be incorporated in the bootstrap block size estimation procedure. Other estimators should be compared and procedures for an adaptive choice of the high level need also to be included in the algorithm. Work on these topics is now in progress.AcknowledgmentsI would like to thank the reviewer for the helpful and valuable comments and suggestions that greatly improved the first version of this paper.References[1] J. Beirlant, Y. Goegebeur, J. Segers, and J. L. Teugels, Statistics of Extremes: Theory and Applications, England, John Wiley & Sons, 2004.]]></page><page Index="74" isMAC="true"><![CDATA[64 M. M. NEVES[2] P. Bu¨hlmann and H. Ku¨nsch, Block length selection in the bootstrap for time series, Comput. Statist. Data Anal. 31, 295–310, 1999.[3] F. Caeiro and M. I. Gomes, On the bootstrap methodology for the estimation of the tail sample fraction, in:Proceedings of COMPSTAT2014, M. Gilli, G. Gonzalles-Rodriguez, A. Nieto-Reyes (eds.), 545–552, 2014.[4] E. Carlstein, The use of subseries methods for estimating the variance of a general statistics from a stationary times series, Ann. Statist. 14, 1171–1179, 1986.[5] J. Danielsson, L. de Haan, L. Peng, and C. G. de Vries, Using a bootstrap method to choose the sample fraction in the tail index estimation, J. Multivariate Anal. 76, 226–248, 2001.[6] G. Draisma, L. de Haan, L. Peng, and T. Pereira, A bootstrap-based method to achieve optimality in estimating the extreme value index, Extremes 2 (4), 367–404, 1999.[7] B. Efron, Bootstrap methods: another look at the jackknife, Ann. Statist. 7, 1–26, 1979.[8] B. Efron, Jackknife-after-Bootstrap Standard Errors and Influence Functions, J. Roy.Statist. Soc. Ser. B 54, 83–111, 1992.[9] M. I. Gomes, Statistical inference in an extremal Markovian Model, in: COMPSTAT–Proc. in Computational Statistics, K. Momirovic, V. Mildner (eds.), Physica-Verlag,Heidelberg, 257–262, 1990.[10] M. I. Gomes, Modelos extremais em esquemas de dependˆencia, in: Estat´ıstica Robusta,Extremos e mais alguns temas, D. Pestana (ed.), I Congresso Ibero-Americano de Es-tad´ıstica e Investigaci´on Operativa, 209–220, Edi¸c˜oes Salamandra, Lisboa, 1992.[11] M. I. Gomes, On the estimation of parameters of rare events in environmental time series, in: Statistics for environment, V. Barnett, K. F. Turkman (eds.), 225–241. JohnWiley & Sons, 1993.[12] M. I. Gomes, F. Caeiro, L. Henriques-Rodrigues, and B. G. Manjunath, Bootstrap meth-ods in statistics of extremes, in: Handbook of Extreme Value Theory and Its Applicationsto Finance and Insurance, F. Longin (ed.), John Wiley & Sons, (in press).[13] M. I. Gomes, F. Figueiredo, and M. M. Neves, Adaptive estimation of heavy right tails:resampling-based methods in action, Extremes 15, 463–489, 2012.[14] M. I. Gomes, F. Figueiredo, M. J. Martins, and M. M. Neves, Resampling methodologiesand reliable tail estimation, South African Statist. J. 49, 1–20, 2015.[15] M. I. Gomes, A. Hall, and C. Miranda, Subsampling techniques and the Jackknife methodology in the estimation of the extremal index, Comput. Statist. Data Anal. 52 (4),2022–2041, 2008.[16] M. I. Gomes, M. J. Martins, and M. M. Neves, Generalised Jackknife-based estimatorsfor univariate extreme-value modeling, Comm. Statist. Theory Methods 42 (7), 1227–1245, 2013.[17] M. I. Gomes, S. Mendonc¸a, and D. Pestana, Adaptive reduced-bias tail index and VaRestimation via the bootstrap methodology, Comm. Statist. Theory Methods 40 (16),2946–2968, 2011.[18] M. I. Gomes and O. Oliveira, The bootstrap methodology in Statistics of Extremes:choice of the optimal sample fraction, Extremes 4 (4), 331–358, 2001.[19] M. I. Gomes and D. Pestana, A sturdy reduced-bias extreme quantile (VaR) estimator,J. Amer. Statist. Assoc. 102 (477), 280–292, 2007.[20] H. L. Gray and W. R. Schucany, The Generalized Jackknife Statistic, Marcel Dekker,1972.]]></page><page Index="75" isMAC="true"><![CDATA[BOOTSTRAP AND JACKKNIFE METHODS IN EXTREMAL INDEX ESTIMATION 65[21] P. Hall, Resampling a coverage pattern, Stochastic Process. Appl. 20, 231–246, 1985.[22] P. Hall, J. L. Horowitz, and B.-Y. Jing, On blocking rules for the bootstrap with depen-dent data, Biometrika 82 (3), 561–574, 1995.[23] T. Hsing, Estimating the parameters of rare events, Stochastic Process. Appl. 37, 117–139, 1991.[24] T. Hsing, Extremal index estimation for a weakly dependent stationary sequence, Ann.Statist. 21, 2043–2071, 1993.[25] J. T. Hsing, J. Hu¨sler, and M. R. Leadbetter, On the exceedance point process for astationary sequence, Probab. Theory Related Fields 78 (1), 97–112, 1988.[26] H. Ku¨nsch, The jackknife and the bootstrap for general stationary observations, Ann.Statist. 17, 1217–1241, 1989.[27] S. Lahiri, On the jackknife after bootstrap method for dependent data and its consistencyproperties, Econometric Theory 18, 79–98, 2002.[28] S. Lahiri, K. Furukawa, and Y-D. Lee, Nonparametric plug-in method for selecting theoptimal block lengths, Stat. Methodol. 4, 292–321, 2007.[29] M. R. Leadbetter, Extremes and local dependence in stationary sequences, Z. Wahrsch.Verw. Gebiete 65 (2), 291–306, 1983.[30] M. R. Leadbetter, G. Lindgren, and H. Rootz´en, Extremes and related properties ofrandom sequences and series, Springer-Verlag, New York, 1983.[31] M. R. Leadbetter and L. Nandagopalan, On exceedance point process for stationarysequences under mild oscillation restrictions, in: Extreme Value Theory, Proceedings, Oberwolfach 1987, J. Hu¨sler, R. D. Reiss (eds.), Lecture Notes in Statist. 52, 69–80, Springer-Verlag, Berlim, 1989.[32] R. Liu and K. Singh, Moving blocks jackknife and bootstrap capture weak dependence, in: Exploring the Limits of Bootstrap, R. Lepage, L. Billard (eds.), 225–248, Wiley, New York, 1992.[33] S. Nandagopalan, Multivariate Extremes and Estimation of the Extremal Index, PhD Thesis, University of North Carolina, Chapel Hill, 1990.[34] M. M. Neves, M. I. Gomes, F. Figueiredo, and D. Prata Gomes, Modeling Extreme Events: Sample Fraction Adaptive Choice in Parameter Estimation, J. Stat.Theory Pract. 9 (1), 184–199, 2014.[35] D. R. Politis and J. P. Romano, A Circular Block-Resampling Procedure for Stationary Data, in: Exploring the Limits of Bootstrap, R. Lepage, L. Billard (eds.), 263–270, Wiley, New York, 1992.[36] D. R. Politis and J. P. Romano, The stationary bootstrap, J. Amer. Statist. Assoc. 89 (428), 1303–1313, 1994.[37] D. R. Politis and H. White, Automatic Block-Length Selection for the Dependent Boot- strap, Econometric Rev. 23 (1), 53–70, 2004.[38] D. Prata Gomes and M. M. Neves, Resampling Methodologies and the Estimation of Parameters of Rare Events, in: Numerical Analysis and Applied Mathematics (ICNAAM 2011), AIP Conf. Proc. 1389, 1475–1478, 2011.[39] D. Prata Gomes and M. M. Neves, Bootstrap and other resampling methodologies in statistics of extremes, Comm. Statist. Simulation Comput. (accepted), 2015a.[40] D. Prata Gomes and M. M. Neves, Adaptive choice and resampling techniques in ex- tremal index estimation, in: Theory and Practice of Risk Assessment, C. Kitsos, T. Oliveira, A. Rigas, S. Gulati (eds.), Springer Proceedings in Mathematics and Statistics (in press), 2015b.]]></page><page Index="76" isMAC="true"><![CDATA[66 M. M. NEVES[41] K. Singh, On the asymptotic accuracy of the Efron’s bootstrap, Ann. Statist. 9, 345–362, 1981.[42] R. Smith and I. Weissman, Estimating the extremal index, J. R. Statist. Soc. B 56, 515–528, 1994.[43] I. Weissman and S. Novak, On blocks and runs estimators of the extremal index, J. Statist. Plann. Inference 66, 281–288, 1998.(M. M. Neves) CEAUL and ISA, Universidade de Lisboa, 1349-017 Tapada da Ajuda, PortugalE-mail address: manela@isa.ulisboa.pt]]></page><page Index="77" isMAC="true"><![CDATA[MEAN SQUARE ERROR IN REGRESSION ESTIMATION WITH FUNCTIONAL DATAPAULO EDUARDO OLIVEIRADedicated to Nazar´e Mendes Lopes on the occasion of her sixtieth birthdayAbstract. The performance of kernel estimators is mainly a↵ected by the choice of the bandwidth parameter. Characterizations for the band- width are usually based on convenient descriptions of the mean square error. So, we prove a characterization of this error criterium, for regres- sion estimation in a functional framework, assuming only the continuity of the regression function. The representation depends on the behaviour of small-ball probabilities, and highlights the influence of the geometry and smoothness of the distribution of the functional variable in the de- scription of this mean square error.1. IntroductionThe problem of approximating regression functions is one of the most classi- cal in statistics and data analysis. Regression estimation is most usually studied in Rd, where the Lebesgue measure plays an essential role, due to its invari- ance to translations. We will be looking at regression estimation considering a functional framework, thus placing ourselves in an infinite dimensional space where no analogue of the Lebesgue measure exists. This functional framework has received a lot of interest from statisticians in recent years. First steps into this direction were made by Ge↵roy [7, 8] who studied the regressogram, later followed by a point process approach by Jacob and Oliveira [10, 11], although not really exploring the functional framework. A really functional ap- proach was used later by Ramsay and Silverman [17] for some case studies,Accepted: 10 April 2015.2010 Mathematics Subject Classification. 62G05, 62G08, 62M09.Key words and phrases. functional data, regression estimation, small-ball probability, mean square error.This work was partially supported by the Centre for Mathematics of the Univer- sity of Coimbra – UID/MAT/00324/2013, funded by the Portuguese Government through FCT/MEC and co-funded by the European Regional Development Fund through the Part- nership Agreement PT2020.67]]></page><page Index="78" isMAC="true"><![CDATA[68 P. E. OLIVEIRAFerraty and Vieu [3, 4], Masry [14] or Ferraty, Mas, and Vieu [6] for regression estimation, followed by many other authors. A good account of essential results of the theory may be found in the monography by Ferraty and Vieu [5].A crucial assumption for the convergence of the estimators concerns the local behaviour of the distribution of the functional variable, known as the small-ball probabilities. Some kind of regularity on these probabilities is always required to prove any convergence result for the kernel estimator. The bandwidth choice, which is well known to be essential for this family of estimators, is usually only indirectly addressed. The classical way to discuss this choice, based on the as- ymptotic description of the mean square error of the regression estimator, has been studied in Ferraty, Mas, and Vieu [6]. Assuming the di↵erentiability of a variant of the regression function and kernels that are compactly supported and bounded away from 0, these authors give a rather compact asymptotic descritp- tion of the mean square error. This characterization is easily used to describe the bandwidth optimal choice. We want to consider kernel functions that are still compactly supported but allow the weights to approach 0 and also avoid di↵erentiability assumptions on the regression function. This choice drives us to make assumptions on an infinite dimensional equivalent of the traditional density functions, giving raise to some integrability issues to be handled. More- over, we may have some knowledge of the distribution of the functional variable. Thus, including this knowledge into the bandwith description leads to a better adapted estimator. With this in mind we give an alternative characterization of the mean square error of the kernel estimator for the regression function. The expression obtained, described in Theorem 4.9 below, is rather long and somewhat intricate, but highlights the role of the geometry and smoothness of the distribution of the functional variable.On the sequel (Xi,Yi), i   1, are equally distributed and independent ran- dom elements, where Yi are random variables, and Xi take values in some normed space S. Our goal is to estimate r(x) = E(Y|X = x), x 2 S. We will denote the conditional second order moment by s(x) = E (Y 2 |X = x), x 2 S. Given a function q : S  ! R, define the local modulus of continuity aswq(x,h)= sup |q(y) q(x)|, x2S,h>0. ky xkhOf course, if q is continuous at x then limh!0 wq (x, h) = 0.2. Assumptions on small-ball probabilitiesWe introduce a first assumption on the behaviour of small-ball probabilities.(SB1) Let F(x,h) = P(kx Xk  h), and assume that, for each x 2 S, limh!0 F (x, h) = 0 and F is partially di↵erentiable with respect to h.]]></page><page Index="79" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 69This is a usual assumption in almost all literature on regression estimation in a functional framework.In order to have some control on the asymptotics we need a more precisedescription of the decrease rate of F when h goes to 0. The control of thisdecrease rate appeared for the first time in Ferraty, Mas, and Vieu [6] throughthe function ⌧0(x, z) = limh!0 F (x,hz) , x 2 S, z 2 [0, 1]. To include in the asyp- F (x,h)totic description of the mean square error some knowledge of the distribution of the functional variable we need some information about the convergence rate towards ⌧0. This is better achieved through a density like variant of ⌧0, that we introduce in the following assumption.(SB2) There exists F0 , defined in S ⇥ [0, 1], such that  ( x , z , h ) = h @ F ( x , z h )   F 0 ( x , z )  ! 0 , w h e n h ! 0 .F(x,h) @hIt is easily seen seen that if ⌧0 is di↵erentiable with respect to the secondvariable then F0(x, z) = @⌧0 (x, z). Thus (SB2) is a translation into density @zlike functions of H3 in [6]. As it will become apparent later (see Example 2.1 (3) below), the two approaches are not equivalent. Indeed, Ferraty, Mas, and Vieu [6] assumption H3, treating distribution functions, regularizes everything thus allowing to apply the Dominated Convergence Theorem, while our den- sity like functions give raise to di culties on handling the asymptotics of the integrals.We will be interested in kernel estimaXtion so, define, for each x 2 S, 1nfb ( x ) = K ( kx   X ik ) , n nF(x,h) hi=1 1 Xngb n ( x ) = Y i K ( k x   X i k ) , nF (x, h) hi=1where K is a real valued function, and h the bandwidth. Further, letgb b ( x ) rbn(x)=n .fn (x)This estimator is the Nadaraya-Watson estimator studied by several authors.Consider the following assumption on the kernel function. (K) K is bounded nonnegative with support [0, 1].Remark that we do not require K Zto be a probability density. Define now (K, x, h) = K(z) (x, z, h) dz. [0,1]]]></page><page Index="80" isMAC="true"><![CDATA[70andf0(x) = ZP. E. OLIVEIRAK(z)F0(x, z) dz, f1(x) = Z[0,1][0,1]The behaviour of the functions   and   depends on the geometry and smooth-ness of the distribution of the random process X. Typically, one expects thatlimh!0  (K, x, h) = 0. Now, this follows from (SB2) if a Dominated Con-vergence Theorem is applicable, but this is not necessarily true, as see inthe examples below. To refer to Ferraty, Mas, and Vieu [6] approach, as-suming oRf course the di↵erentiability of ⌧0 and K, it is easily verified thatf0(x) = 1 K0(z)@⌧0 (x,z)dz. Moreover, in [6] it is assumed, besides the distri-0 @zbution function version of (SB2), that the kernel satisfied (K) and also thatit is decreasing with K(1) > 0. Thus, kernels with weights decreasing smoothly to 0, such as K(z) = 1   z✓, ✓ > 0, are not allowed under the assumptions in [6]. The need for such extra asumptions on K is due to the peculiarities of the control of the asymptotics, made through distribution functions and not needing the convergence rate towards ⌧0. However, kernel functions satisfying the assumptions of [6] leave us with an estimator the resambles the regresso- gram, as there is a sudden cut-o↵ on the weights once the observation becames to far from the reference point x.Examples 2.1. We shall next describe in more detail the functions   and   in a few significative examples. The models (1)–(3) are also discussed in Ferraty, Mas, and Vieu [6] describing the characterizations adapted to their assumptions.(1) Let F(x,h) = f(x)h ,   > 0. If   2 N this model includes the case of X being a  -dimensional random vector. For general   > 0, the model corresponds to fractal type processes of order   (see Ferraty and Vieu [3, 5]). It is easily verified that F0(x, z) =  z  1,  (x, z, h) = 0, thus, also  (K, x, h) = 0.(2) With respect to the previous example, we may relax the concentrationrate of the probability mass around x assuming F (x,h) = f (x)h  |log h|k , , k > 0. Then, F0(x,z) =  z  1,  (x,z,h) = kz  1 , andK2(z)F0(z) dz.|log h| kZ1 (K, x, h) = |log h| K(z)z  1 dz  ! 0, 0if this last integral is finite. Thus  (K,x,h) does converge to 0, butwith a rather slow rate.(3) Assume F(x,h) = f(x)h e h k,  , k > 0. This model for the distri-bution of X includes di↵usion processes, fractional Brownian motions, fractional Brownian sheets or fractional Ornstein-Uhlenbeck processes.]]></page><page Index="81" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 71 It is easily verified that F0(x,z) = 0 if z 6= 1, and F0(x,1) = +1.Thus,  (x,z,h) =   + k  z  1exp h k  (zh) k  if z 6= 1, and zkhk (x,1,h) =  1. Remark that the Dominated Convergence Theorem does not apply, as the functions  (x,·,h) are not uniformly bounded with respect to h. Now, if we assume K is di↵erentiable, it follows, using integration by parts, thatZ1 0   1   k  k   (K,x,h)=K(1)  K (z)z exp h  (zh) dz.0The popular choice K(z) = I[0,1](z) leads to  (K,x,h) = 1, thus not converging to 0. However, this di culty may be overcome by choosing K(z) = 1   z✓, withZ ✓   max(0, 2   (  + k)), as in such case11   k  k  k  (K,x,h)✓ zk+1 exp h  (zh) dz=✓h .0That is, for this model we have a furthter argument to be interested inkernels such that the weights approach 0.(4) Assume S = L2([0,1]d) and X is a Gaussian process on S,  (u,v) =Cov(X(u), X(v)) isRthe auto-covariance function of X, and consider the operator ⇤f(u)= [0,1]d  (u,v)f(v) d(dv),f 2Sand d theLebesguemeasure on [0, 1]d. If x belongs to the reproducing Hilbert space inducedby   then P(kx Xk  h) = P(kXk  h) (see Li and Shao [13]), sowe may shift the small-ball problem to the origin. Assume the operator ⇤ haseigenvalues n,n 1,and#{n: n >t}⇠'(1),where' tis a slowly varying function. In this case, it has been shoRwn recently by Karol and Nazarov [12] that P(kXk  h) ⇠ exp⇣ 1 u '(z) dz⌘,21zwhere '(u) ⇠ h. Assume that  n ⇠ n ↵, ↵ > 1 (the Brownian motion2ucorresponds to ↵ = 2). Then '(z) = z1/↵ and it is easily verified that1 ↵ 1 ↵ P(kXk  h) ⇠ e↵/2 exp⇣ 22↵ 1 ↵h ↵ ⌘,which corresponds to model (3) with   = 0 and k = ↵ . ↵ 13. Convergence of the estimatorUsing Lemma 2.1 in [10] it follows that, for " > 0 small enough, n  b o    rb n ( x )   E gb n ( x )     > "n E f n ( x ) o n         ( E fb ( x ) ) 2 o(3.1)⇢ |gb(x) Egb(x)|>"Efb(x) [  fb(x) Efb(x) >" n . n n 4n n n 4Egbn(x)]]></page><page Index="82" isMAC="true"><![CDATA[72 P. E. OLIVEIRAWe start by describing the asymptotic behaviour of the expectation and variance of fb (x) and gb (x).nnLemma 3.1. Assume that (SB1), (SB2) and (K) hold. ThenEfb(x)=f (x)+ (K,x,h), n0◆ .Proof. The proof is straightforward after writing the mathematical expecta-Var fb (x) = f1(x) +  (K2, x, h) + o ✓ 1nnF (x, h) nF (x, h)tions as integrals over [0, 1]. ⇤Note that if lim  (K,x,h) = 0 the behaviour of the variance has a h!0b f1(x) ⇣1⌘simpler description: Var fn(x) = nF (x,h) +o nF (x,h) . Similar comments apply throughout this paper but, to avoid repetitions, we will not be referring to suchvariations at each place.Lemma 3.2. Assume that (SB1), (SB2) and (K) hold. Then E gb n ( x ) = ( r ( x ) + w r ( x , h ) ) ( f 0 ( x ) +   ( K , x , h ) ) ,(s(x) + ws(x, h))(f1(x) +  (K2, x, h)) ✓ 1 ◆ Vargbn(x)= Z nF(x,h) +o nF(x,h) .Proof. The proof follows as for the previous lemma, by writingE gb n ( x ) = 1 ( r ( x ) + r ( u )   r ( x ) ) K ( k x   u k ) P X ( d u )F(x,h) h  1 Z (r(x) r(u))K(kx uk)P (du) w (x,h)(f (x)+ (K,x,h)).and remarking that F(x,h) hX r 0⇤The following result is now obvious.Corollary 3.3. Assume that wr(x,·), ws(x,·),  (K,x.·) and  (K2,x,·) arebounded, (SB1), (SB2), (K) hold andnF (x, h)  ! +1. (3.2)Then Varfb (x)  ! 0 and Vargb (x)  ! 0. nn]]></page><page Index="83" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 73Theorem 3.4. Let r be continuous and assume that (SB1), (SB2) and (K) hold. Then rbn(x) is asymptotically unbiased.Proof. Indeed, let ren(x) = E gbn(x) . It follows from the above characterizations bEfn(x)that ren(x) = r(x) + wr(x, h)  ! r(x), as r is continuous. ⇤We now state convergence results about the estimator. The proofs are straight- forward adaptations of the proof of Theorems 3.2, 3.3 and 3.4 in Oliveira [16], so we will not include them here.Theorem 3.5. Assume s is continuous, (SB1), (SB2) and (K) hold. If there exists M1 >0 such that, for every ` 2, E(Y`|X =x)M1``!s(x) andnF (x, h)  ! +1, (3.3) log nthen rbn(x)   E rbn(x) converges almost completely to zero with a convergence rate of order ⇣ log n ⌘1/2.nF (x,h)Theorem 3.6. Assume r is continuous, (SB1), (SB2) and (K) hold. Thenp 0 f1(x)+ lim  (K2, x, h) 1 d @ h!0 2 AnF(x,h)(rbn(x) Erbn(x)) !N 0, f0(x)+lim  (K,x,h) (s(x) r (x)) . h!04. The mean square errorWe want to characterize the asymptotic behaviour of E (rbn(x)   r(x))2. For this, as done in Bosq and Cheze [2] and later explored by Oliveira [15] and Bensa¨ıd and Fabre [1], we consider the decompositionE(rb (x) r(x))2 = r2(x)E(fb (x) f (x))2 n f02(x)n0+ 1 E ( gb n b( x )   r ( x ) f 0 ( x ) ) 2 f 02 ( x )  2r(x) E ⇣(fn(x)   f0(x))(gbn(x)   r(x)f0(x))⌘ f 02 ( x )(4.1)⇣2 2 bb 2⌘ f02(x)n n01+ E (rb (x) r (x))(f (x) f (x))  2 E ⇣(rbn(x)   r(x))(fn(x)   f0(x))(gbn(x)   r(x)f(x))⌘ . f 02 ( x )]]></page><page Index="84" isMAC="true"><![CDATA[74 P. E. OLIVEIRAWe will now go through each term above, characterizing their asymptotics. The first result describes the mean square error of fb (x) and gb (x), handling thetwo first terms. It is an immediate consequence of Lemmas 3.1 and 3.2.Lemma 4.1. Assume that (SB1), (SB2) and (K) hold. ThenandnF (x, h)E (gbn(x)   r(x)f0(x))2 = (s(x) + ws(x, h))(f1(x) +  (K2, x, h))E ( fb ( x )   f ( x ) ) 2 = f 1 +   ( K 2 , x , h ) +   2 ( K , x , h ) .n0nF (x, h) +r(x) 2(K, x, h) + wr2(x, h)(f0(x) +  (K, x, h))2+2r(x) (K,x,h)wr(x,h)(f0(x)+ (K,x,h))+o✓ 1 ◆. nF (x, h)nnLemma 4.2. Let r be continuous and assume that (SB1), (SB2) and (K) hold. ThenE ⇣(fb (x)   f (x))(gb (x)   r(x)f (x))⌘ n0n0= r(x) (f1(x) +  (K2, x, h)) + r(x) 2(K, x, h) nF (x, h)+wr(x,h)(f0(x)+ (K,x,h)) (K,x,h)+o✓ 1 ◆. nF (x, h)term follows from Lemmas 3.1 and 3.2. On the other handProbof. Write E  (fb (x)   f (x))(gb (x)   r(x)f (x))  = Cov(fb (x), gb (x)) + n0n0nn b    E (fn(x)   f0(x)) E (gbn(x)   r(x)f0(x)) . The characterization of the lastC o v ( f n ( x ) , gb n ( x ) ) ✓= 1 E YK(kx Xk) +o 1◆= 1 E (r(x)+r(X) r(x))K(kx Xk) +o✓ 1 ◆n2F 2(x, h) h nF (x, h)n2F 2(x, h) h nF (x, h)= r(x)  f1(x)+ (K2,x,h) +o✓ 1 ◆. nF (x, h) nF (x, h)The conclusion now follows imediately. ⇤]]></page><page Index="85" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 75We will need the characterization of the asymptotics of E (fb (x)   f (x))4 . n0Lemma 4.3. Assume that wr(x,·), ws(x,·),  (K,x.·) and  (K2,x,·) are bounded, and (SB1), (SB2), (K) hold. ThenE(fb(x) f (x))4 =O✓ 1 + 2(K,x,h)+ 4(K,x,h)◆.n0Proof. Write ⇣ ⌘4n2F 2(x, h) nF (x, h) E(fb(x) f (x))4 =E (fb(x) Efb(x))+(Efb(x) f (x))n0nnn0=E(fb(x) Efb(x))4 +4E(fb(x) Efb(x))3(Efb(x) f (x)) nnnnn0+6E(fb(x) Efb(x))2(Efb(x) f (x))2 +(Efb(x) f (x))4. nnn0n0We shall now characterize each term of the previous expansion. Obviously (Efb(x) f (x))4 = 4(K,x,h).n0 bbb2b2E(fn(x) Efn(x)) (Efn(x) f0(x)) =Varfn(x) 2(K,x,h)=O✓ 2(K,x,h) + 4(K,x,h)◆.nF (x, h)For the second term in the expansion, we have, expanding again and takinginto account that the terms are centered and independent,E(fb (x) Efb (x))3 = 1 E ⇣K(kx Xk) EK(kx Xk)⌘3n n n2F3(x,h) h h= 1 ⇣EK3(kx Xk) 3EK2(kx Xk)EK(kx Xk)+2(EK(kx Xk))3⌘n2F3(x,h) h h h h =O✓ 1 ◆.n2F 2(x, h)Finally, expanding once more, we find that E (fb (x) E fb (x))4 = O  1  .n n n2F2(x,h) Summing up all the characterizations, the result follows.⇤To continue with the analysis of the terms in (4.1), write first⇣2 2 b 2⌘ E (rb (x) r (x))(f (x) f (x))nn0⇣2 2 b 2⌘2 2 b 2=E (rb (x) re (x))(f (x) f (x)) +(re (x) r (x))E(f (x) f (x)) . nnn0n n0]]></page><page Index="86" isMAC="true"><![CDATA[76 P. E. OLIVEIRAAs re (x)  ! r(x), taking into account Lemma 4.1, the second term abovenF (x,h)numbers and write⇣2 2 b 2⌘ E (rb (x) re (x))(f (x) f (x))⇣⌘is easily seen to be an o 1 + 2(K,x,h) . Now, let ↵n and "n be realnnnn0 ⇣22b2⌘= E (rb (x) re (x))(f (x) f (x)) I⇣ n n n 0 {rbn (x)↵n ,|rbn (x) ren (x)|"n } ⌘22b2 +E (rb (x) re (x))(f (x) f (x)) Innn0 {rbn (x)↵n ,|rbn (x) ren (x)|>"n } ⇣22b2⌘+E (rb (x) re (x))(f (x) f (x)) I =:A +B +C . n n n 0 {rbn(x)>↵n} nnnLemma 4.4. Assume that (SB1), (SB2) and (K) hold. If ↵n % +1, "n & 0 are such that ↵n"n  ! 0 thenAn =o✓ 1 + 2(K,x,h)◆. nF (x, h)Proof. The result is obvious by noting that A  (↵ + re (x))" E (fb (x)   nnnnnf0(x))2 and taking into account Lemma 4.1. ⇤ To bound Bn, we start by remarking thatB (↵ +re (x))2⇣E(fb(x) f (x))4⌘1/2(P(|rb (x) re (x)|>" ))1/2, nnnn0nnnand using (3.1) to obtain P(|rbn(x)   ren(x)| > "n)⇣ b ⌘ ⇣     b b     P |gbn(x) Egbn(x)|>"nEfn(x) +P  fn(x) Efn(x) >"n( E fb ( x ) ) 2 ⌘ n .◆,4 4E gb n ( x )Using now Markov’s inequality it follows that P⇣|gb(x) Egb(x)|>"nEfb(x)⌘ 16Vargbn(x) = 1O✓ 1n n 4 n " 2 ( E fb ( x ) ) 2 " 2n n F ( x , h ) nnand analogously for the other probability term.Lemma 4.5. Assume that (SB1), (SB2), (K) and (3.2) hold. ThenBn =o✓ 1 +  (K,x,h) + 2(K,x,h)◆. nF (x, h) n1/2 F 1/2 (x, h)]]></page><page Index="87" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 77Proof. Taking into account the discussion above, it is enough to control ↵2 (E (fb (x)   f (x))4)1/2(P(|gb (x)   E gb (x)| > " ))1/2nn0nnn ↵n2✓1 (K,x,h)2 ◆✓1◆= " O nF(x,h)+n1/2F1/2(x,h)+  (K,x,h) O n1/2F1/2(x,h) . nNowchoose↵, >0suchthat2↵+ <1 andput↵n=(nF(x,h))↵and"n=2= (nF(x,h))2↵+  1/2  ! 0, so the termLemma 4.7. Assume that wr(x,·), ws(x,·),  (K,x.·) and  (K2,x,·) are bounded, (SB1), (SB2), (K) and (3.2) hold. If, for some ↵ > 0,(nF(x,h))  . Then↵2n"n n1/2 F 1/2 (x,h)⌘+  2(K, x, h) . The term corresponding⌘⇣to P  fb (x)   E fb (x)  > (E fb (x))2above is an o1+  (K,x,h) nF (x,h)   n1/2 F 1/2 (x,h)⇣ is treated analogously. ⇤ Remark 4.6. If we choose   > ↵ then ↵n"n  ! 0."n nn n 4 Egbn(x)then4E(rb(x)I n↵ ) !0, { rb n ( x ) > ( n F ( x , h ) ) }+  (K,x,h) + 2(K,x,h)◆. n1/2 F 1/2 (x, h)(4.2)Cn =o✓Proof. Put ↵n = (nF (x, h))↵  ! +1 and write1nF (x, h)⇣bb⌘222Cn = E (rbn(x)   ren(x)(x))(fn(x)   f0(x)) I{rbn(x)>↵n}⇣ E (fn (x)   f0 (x)) E (rbn (x)I{rbn (x)>↵n } )b 4⌘1/2   4  1/2 ⇣ 4⌘1/2 2 1/2E (fn(x)   f0(x)) ren(x) (P(rbn(x) > ↵n)).⇤+The result is now obvious taking into account (4.2).To complete the characterization of the mean square error we still have to control the final term in (4.1). The asymptotic behaviour of this last term depends more directly on the geometry and smoothness the distribution of X.]]></page><page Index="88" isMAC="true"><![CDATA[78 P. E. OLIVEIRALemma 4.8. Assume wr(x,·), ws(x,·),  (K,x.·) and  (K2,x,·) are bounded,(SB1), (SB2), (K), (3.2) and (4.2), for some ↵ 2 (0, 1 ), hold. Then 6E ⇣(rb (x)   r(x))(fb (x)   f (x))(gb (x)   r(x)f(x))⌘ nn0n=o 1 +  1/2(K,x,h) + (K,x,h)+ 1/2(K,x,h)+wr(x,h) nF (x, h) n3/4F 3/4(x, h) n1/2F 1/2(x, h)+ 3/2(K,x,h)+ 1/2(K,x,h)wr(x,h)+ (K,x,h) n1/4 F 1/4 (x, h)+ 2(K,x,h)+ 3/2(K,x,h)+ (K,x,h)wr(x,h)!. Proof. Using Cauchy’s inequality we haveE (rbn(x)   r(x))(fn(x)   f0(x))(gbn(x)   r(x)f(x)) ⇣bb⌘ ⇣ E ( rb n ( x )   r ( x ) ) 2 ( f n ( x )   f 0 ( x ) ) 2 ⌘ 1 / 2 ⇣ E ( gb n ( x )   r ( x ) f 0 ( x ) ) 2 ⌘ 1 / 2 . The last factor above has been characterized in Lemma 4.1. As what regardsthe first factor, its square is bounded above by⇣ 2bb2⌘ 2E (rbn(x)   ren(x)) (fn(x)   f0(x))(4.3) +2(ren(x)   r(x))2E (fn(x)   f0(x))2.As re (x)  ! r(x) and taking into account Lemma 4.1, the second term is an ⇣n⌘o 1 + 2(K,x,h) .Choose↵n =(nF(x,h))↵,"n =(nF(x,h))   where nF (x,h) >↵>0and2↵+ < 1.Notethat,as↵< 1,thischoiceisalwayspossible. 26Decompositing as before, the first term in (4.3) is bounded above byh⇣2b2⌘ (↵n + ren(x)) E (rbn(x)   ren(x))b(fn(x)   f0(x)) I{rbn(x)↵n,|rbn(x) ren(x)|"n}+E ⇣(rb (x)   re (x))2(fb (x)   f (x))2I ⌘in n n0 {rbn (x)↵n ,|rbn (x) ren (x)|>"n } +E ⇣(rbn (x)   ren (x))2 (fn (x)   f0 (x))2 I{rbn (x)>↵n } ⌘ .]]></page><page Index="89" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 79The first and second terms of this decomposition are controled as in Lemmas 4.4⇣ 1 b 2  (K,x,h) ⌘and 4.5 to obtain an o nF(x,h) +  (K,x,h)+ n1/2F1/2(x,h) . FinallyE ⇣(rbn(x)   ren(x))2(fn(x)  bf0(x))2I{rbn(x)>↵n}⌘ ⇣222⌘2E (rb (x)+re (x))(f (x) f (x)) In n n 0 { rb n ( x ) > ↵ n }= o ✓ 1 +  2(K, x, h) +  (K, x, h) ◆ , nF (x, h) n1/2 F 1/2 (x, h)taking into account (4.2). Collecting these upper bounds, the result follows. ⇤ To describe the mean square error for the regression estimator it remains toput together the characterizations just obtained for each term of (4.1). Theorem 4.9. Assume that r is continuous, ws(x, ·),  (K, x.·) and  (K2, x, ·)are bounded, (SB1), (SB2), (K), (3.2) and (4.2), for some ↵ 2 (0, 1 ), hold.Then,E(rbn(x) r(x))2= f1(x) +  (K2, x, h) ✓ r2(x) + s(x) + ws(x)   2r(x) ◆6nF (x, h) f02 (x) f02 (x) f02 (x) +wr2(x,h)(f0(x)+ (K,x,h))2 2r(x)wr(x,h) (K,x,h)(f0(x)+ (K,x,h))+of02 (x) f02 (x)1 +  1/2(K,x,h) +  (K,x,h)+ 1/2(K,x,h)+wr(x,h)nF (x, h) n3/4F 3/4(x, h) n1/2F 1/2(x, h) + 3/2(K,x,h)+ 1/2(K,x,h)wr(x,h)+ (K,x,h)n1/4 F 1/4 (x, h) + 2(K,x,h)+ 3/2(K,x,h)+ (K,x,h)wr(x,h)!.Let us get back to the case where limh!0  (K, x, h) = 0. We may trace the calculations above to find that, under this additional assumption and the]]></page><page Index="90" isMAC="true"><![CDATA[80 P. E. OLIVEIRAcontinuity of the second conditional moment s, we have E(rbn(x) r(x))2= f1(x) ✓ r2(x) + s(x) + ws(x)   2r(x) ◆ + wr2(x, h) nF (x, h) f02 (x) f02 (x) f02 (x) f0 (x)+o✓ 1 + wr(x,h) + (K,x,h)◆. n1/4F 1/4(x, h)  1/2(K, x, h)n1/4F 1/4(x, h)References[1] N. Bensa¨ıd and J. P. Fabre, Optimal asymptotic quadratic error of kernel estimators of Radon-Nikodym derivatives for strong mixing data, J. Nonparametr. Stat. 19, 77–88, 2007.[2] D. Bosq and N. Cheze, Erreur quadratique asymptotique optimale de l’estimateur non param´etrique de la r´egression pour des observations discretis´ees d’un processus station- naire `a temps continu, C. R. Acad. Sci. Paris S´er. I Math. 317 (9), 891–894, 1993.[3] F. Ferraty and P. Vieu, Dimension fractale et estimation de la r´egression dans des espaces vectoriels semi-norm´es, C. R. Acad. Sci. Paris S´er. I Math. 330, 139–142, 2000.[4] F. Ferraty and P. Vieu, Nonparametric models for functional data, with application in regression, time series prediction and curve estimation, The International Conference on Recent Trends and Directions in Nonparametric Statistics, J. Nonparametr. Stat. 16, 111–125, 2004.[5] F. Ferraty and P. Vieu, Nonparametric Functional Data Analysis. Theory and Practice, New York, Springer, 2006.[6] F.Ferraty,A.Mas,andP.Vieu,NonparametricRegressiononFunctionalData:Inference and practical aspects, Aust. N. Z. J. Stat. 49, 267–286, 2007.[7] J. Ge↵roy, Sur l’estimation d’une densit´e dans un espace m´etrique, C. R. Acad. Sci. Paris S´er. A 278, 1449–1452, 1974.[8] J. Ge↵roy, Etude de la convergence du r´egressogramme, Publ. Inst. Statist. Univ. Paris 25 (1-2), 41–56, 1980.[9] P. Jacob and L. Ni´er´e, Contribution a` l’estimation des lois de Palm d’une mesure al´eatoire, Publ. Inst. Statist. Univ. Paris 35 (2), 39–49, 1990.[10] P. Jacob and P. E. Oliveira, A general approach to nonparametric histogram estimation, Statistics 27, 73–92, 1995.[11] P. Jacob and P. E. Oliveira, Kernel estimators of general Radon-Nikodym derivatives, Statistics 30, 25–46, 1997.[12] A. Karol and A. Nazarov, Small ball probabilities for smooth Gaussian fields and tensor products of compact operators, Math. Nachr. 287, 595–609, 2014.[13] W. V. Li and Q. M. Shao, Gaussian processes: inequalities, small-ball probabilities and applications, in: Stochastic Processes: Theory And Methods, C.R. Rao, D. Shanbhag (eds.), Handbook of Statistics 19, 533–597, Elsevier, 2001.[14] E. Masry, Nonparametric regression estimation for dependent functional data: asymp- totic normality, Stochastic Process. Appl. 115, 155–177, 2005.]]></page><page Index="91" isMAC="true"><![CDATA[REGRESSION ESTIMATION FOR FUNCTIONAL DATA 81[15] P. E. Oliveira, Mean square error for histograms when estimating Radon-Nikodym derivatives, Portugal. Math. 57, 1–16, 2000.[16] P. E. Oliveira, Nonparametric density and regression estimation for functional data, Preprint, Publ. Dep. Matema´tica Univ. Coimbra 05-09, 2005.[17] J. Ramsay and B. Silverman, Applied Functional Data Analysis: Methods and Case Studies, New York, Springer, 2002.(P. E. Oliveira) CMUC, Department of Mathematics, University of Coimbra, 3001-501 Coimbra, PortugalE-mail address: paulo@mat.uc.pt]]></page><page Index="92" isMAC="true"><![CDATA[]]></page><page Index="93" isMAC="true"><![CDATA[INTEGER-VALUED SELF-EXCITING PERIODIC THRESHOLD AUTOREGRESSIVE PROCESSESISABEL PEREIRA, MANUEL SCOTTO, AND RAQUEL NICOLETTEDedicated to Maria de Nazar´e LopesAbstract. In this paper, the periodic self-exciting threshold integer- valued autoregressive model of order one with period T driven by a periodic sequence of independent Poisson-distributed random variables is introduced and analyzed in detail. Basic probabilistic and statistical properties of the model are discussed as well as parameter estimation and forecasting.1. IntroductionModeling the temporal dependence and evolution of integer-valued (and in particular low counts) time series is an area of research which is gaining im- portance in time series analysis. The problem of developing models for integer- valued time series is, indeed, very challenging because traditional approaches based on Gaussian autoregressive-moving average processes, are of little use to accurately describe time series defined over finite range of counts or exhibiting features such low counts, over dispersion, asymmetric marginal distributions, or excess of zeros. The need to analyze such data adequately led to a multi- plicity of approaches and a diversification of models that explicitly account for such features.Recently, models for dealing with integer-valued time series exhibiting the so-called piecewise phenomenon have been proposed in the literature. The fun- damental reason for introducing such class of models is the need to model ran- dom cyclic behavior that exists in many time series. In the continuous-valuedAccepted: 14 March 2015.2010 Mathematics Subject Classification. 62M10; 91B70; 60G10 .Key words and phrases. Count processes, Binomial thinning, Threshold models.This work was supported by Portuguese funds through the CIDMA - Center for Re- search and Development in Mathematics and Applications, and the Portuguese Foundation for Science and Technology (FCT Fundac¸˜ao para a Ciˆencia e a Tecnologia), within project UID/MAT/04106/2013.83]]></page><page Index="94" isMAC="true"><![CDATA[84 I. PEREIRA, M. SCOTTO, AND R. NICOLETTEcase, threshold models are typically characterized by having a linear (ARMA) structure in each regime; see, e.g., Turkman et al. (2014) for details. However, in the field of integer-valued time series modelling little research has been done so far to develop models to cope with time series of counts exhibiting piecewise- type patterns. For this purpose, Monteiro et al. (2012) introduced a class of self-exciting threshold integer-valued autoregressive (SETINAR, in short) models of order one and two regimes, defined by the recursive equationXt =⇢↵1 Xt 1+Zt,Xt 1R. ↵2  Xt 1 +Zt, Xt 1 >RHere, (Zt) constitutes a sequence of integer-valued random variables (r.v’s), and R represents the threshold level. The “↵i ” is the binoPmial thinning operator of Steutel and van Harn (1979). It is defined as ↵ X := Xi=1 Yi, for X with rangeN0 = {0, 1, . . .}, where the(i.i.d.) Bernoulli variables with probability ↵ 2 (0; 1). The authors discussed probabilistic and statistical properties related with this class of models. Note that the SETINAR models fall within the state-dependent thinning class.In this article, we introduce the periodic self-exciting threshold integer- valued autoregressive model of order one with two regimes (hereafter referred to as PSETINAR(2; 1, 1)T ) which generalizes the SETINAR model by considering periodically varying threshold levels. For this class of models, we investigate some basic probabilistic and statistical properties. Furthermore, parameter es- timation and forecasting are also addressed. Finally some concluding remarks are given.The PSETINAR(2; 1, 1)T model is defined through the recursive equationYi’s are independent and identically distributed( ↵(1)   X + Z(1), X  RXt = j t 1 t t d t , t 2 N0, (1.1)↵(2)   X + Z(2), X > R j t 1 t t d twithRt =rj,fort=j+sT,j=1,...,Tands2N0.Notethatforthe jth-period we haveXj+sT = (↵(1)  Xj+sT  1 + Z(1) )I(1) + (↵(2)  Xj+sT  1 + Z(2) )I(2), j j+sT j j j+sT jwith I(k), for k = {1, 2⇢}, defined as jI(1) := 1, Xj+sT d rj , I(2) =1 I(1). j 0, Xj+sT d >rj j j(1.2)(1.3)The threshold parameter Rt (which is assumed to be known) represents the level of the process and the regime switch is triggered by the lag-d value of the series.]]></page><page Index="95" isMAC="true"><![CDATA[PERIODIC THRESHOLD PROCESSES 85Note that the model in (1.1) can be represented asXt =  t   Xt 1 + Zt, (1.4)where Zt = Z(1)I(1) + Z(2)I(2),  t = ↵j ⌘ ↵(1)I(1) + ↵(2)I(2), such as ↵j 2 jj jj jj jj(0, 1), t = j + sT , j = 1, . . . , T , and s 2 N. Furthermore, the thinning operator   is defined asandBMt :=B Zt =B T 1 2.... ... ... . C B Z2+tT CXt 1 Xt 1 t   Xt 1 =d X Ui,t(↵(1))I(1) + X Ui,t(↵(2))I(2)B Q ↵T 1 i Bi=0@ T 1 1T 1 3 C B@ . CA. Q ↵T 1 i ... 1 0C .Q ↵T i i=0Q ↵T i ... ↵T 1 i=0jj jj i=1 i=1with (Ui,t(↵(1))) and (Ui,t(↵(2))), i 2 N, being periodic sequences of i.i.d. jjBernoulli r.v’s with success probabilities P(Ui,t(↵(k)) = 1) = ↵(k), for k 2 jj{1,2}. Moreover, the innovation process (Zt) forms a periodic sequence of in- dependent Poisson-distributed r.v’s with mean vt, Zt ⇠ P o(vt), where vt =  j , t = j +sT, j = 1,...,T, s 2 N0. It is assumed that Zt is independent of Xt 1 and ↵t   Xt 1, for every t.The PSETINAR(2; 1, 1)T process in (1.1) can be embedded in the following vectorial formYt = A   Yt 1 + Mt,being Yt = [X1+tT X2+tT · · · XT +tT ]0, where 0 denotes matrix transpose,(1.5)000... ↵1 1 B 0 0 ... ↵1↵2 CA=B . ... ...B@ Q CA. C 00... ↵T jT 1 j=0010...001B ↵2 1 ...00C0 1B ↵3↵2↵3 ... 0 0 C Z1+tTi=0 C ZT+tT T 1 2 A]]></page><page Index="96" isMAC="true"><![CDATA[86 I. PEREIRA, M. SCOTTO, AND R. NICOLETTENote that, A   Y is a T -dimensional random vector with i-th componentT 1 0i1[A Y]i = X0 Xj+tT +@Y↵jA XT+tT,j=1 j=1for i = 1, . . . , T . The components of B   Z can be defined similarly.The rest of the paper is organized as follows: In Section 2, we demonstrate the existence of a strictly ciclostationary PSETINAR(2; 1, 1)T process satisfying (1.5). Furthermore, the expression for the periodic mean is given. Parameter estimation is covered in Section 3. Finally, forecasting is discussed in Section 4.2. Some properties of the PSETINAR model with two regimesLet (Xt) be the PSETINAR(2; 1, 1)T process defined in (1.2). We first provethatthereexistsastrictlyciclostationaryPSETINAR(2;1,1)T processsatisfying(1.2). Note that, since ↵(k) 2 (0,1) for j = 1,...,T and k = 1,2, and that jP(Z(k) = 0) 2 (0,1), for k = 1,2, it follows by Lemma 3 in Franke and j +sTSubba Rao (1995) that any solution (Yt) of (1.5) is an irreducible and aperiodic Markov chain on N0. Thus, the existence of a ciclostationary solution of (1.5) relies upon the largest eigenvalue of the A matrix in (1.5). The result is quoted below.Proposition 2.1. Let (Yt) be the PSETINAR(2; 1, 1)T process defined in (1.5). If E||Mt|| < +1 and if the largest eigenvalue, say ⌘, of A is less than one, then there exists a strictly ciclostationary PSETINAR(2; 1, 1)T process satisfying (1.5).Proof. Proposition B in Dion et al. (1995, p. 126) allows us to conclude that ↵(1)I(1) + ↵(2)I(2) < 1, j = 1, . . . , T , ⌘ < 1. (2.1)Conditions in (2.1) imply that all roots of the characteristic polynomial of A lie inside the unit circle. Furthermore, if E||Zt|| < 1 it follows by Theorem 1 in Franke and Subba Rao (1995) that there exists a strictly ciclostationary PSETINAR(2; 1, 1)T process satisfying (1.5). ⇤Without employing any distributional assumption on the periodic sequences Z(1) and Z(2), the periodic mean of the process is given in the next result. Forttjj jjsimplicity in notation we defineu(k1,k2,...,kj) :=EhXsT|X1+sT d 2r(k1),X2+sT d 2r(k2),...,Xj+sT d 2r(kj)i, 1,2,...,j 1 2 jp(k1,k2,...,kj) :=PhX1+sT d 2r(k1),X2+sT d 2r(k2),...,Xj+sT d 2r(kj)i, 1,2,...,j 1 2 j]]></page><page Index="97" isMAC="true"><![CDATA[PERIODIC THRESHOLD PROCESSES 87= {1, 2}(, where rj denotes the regime corresponding to thethe process takes the formfor k1, k2, . . . , kj period j, i.e.,, t2N0. (2.2) Lemma 2.2. Let (Xt ) be the PSETINAR(2; 1, 1)T process in (1.5). The mean ofr(1), Xj+sT d rj rj = jr(2), Xj+sT d >rj jP2 P2 P2 h (k1) (k2) (kj) (k1,k2,...,kj) ! (k1,k2,...,kj)i E[Xt] = ··· ↵1 ↵2 ···↵j ⇥u1,2,...,j ⇥p1,2,...,jfort=j+sT,j=1,...,T ands2N0.3. Parameters estimationIn this section, we consider the parameter estimation of the PSETINAR(2; 1, 1)T process. In particular, the conditional least squares (CLS) and conditional max- imum likelihood methods are adopted. For this purpose, let (X1, . . . , Xn) be a sequence of r.v’s satisfying (1.4) and denote by✓ := (↵(1),↵(2),  ,...,↵(1),↵(2),  ), 111TTTthe vector of unkown parameters. Recall that Rt is assumed to be known. 3.1. Conditional least squares estimators. The CLS-estimators,↵ˆ(2) , ˆT,CLS), T,CLSk1 =1k2 =1 kj =1Pj P2 P2 ( k l + 1 ) ( k j ) k l + 1 , . . . , k j+  l ··· ↵l+1 ···↵j ⇥pl+1,...,j , l=1 kl+1 =1 kj =1✓ˆCLS := (↵ˆ(1) ,↵ˆ(2) , ˆ1,CLS,...,↵ˆ(1)1,CLS 1,CLSare obtained by minimizing the expressionT,CLS◆2with N and T denoting the number of complete cycles and number of periods,N 1 T ✓Q(✓):= XX Xj+sT  gj(✓j,Xj+sT 1)s=0 j=1respectively. Moreover, ✓j := ⇣↵(1),↵(2), j⌘ and the function gj takes the form jjgj(✓j,Xj+sT 1)=↵(1)Xj+sT 1I(1) +↵(2)Xj+sT 1I(2) + j. jjjjSolving the systems of 8the form @Q> > > >: ( 1 ) = 0 >< @ ↵ j> @↵j@Q = 0, (2)j = 1,...,T,(3.1)@Q =0 @ j]]></page><page Index="98" isMAC="true"><![CDATA[88 I. PEREIRA, M. SCOTTO, AND R. NICOLETTEwe obtain the following set of CLS-estimators8>>> > ↵ˆ ( 1 )> j,MQC> j,MQC P=◆2 Xj+sT  1IjNXj+sT Xj+sT  1Ij   Xj+sTs = 0 s = 0 s = 0N 1 (1) N 1 N 1 PPP(1) Xj+sT  1Ij> >N 1 ✓N 1N P X2 I(1)   P Xj+sT  1I(1)s=0 j+sT 1 j s=0 j N 1 (2) N 1 N 1><P PP > N Xj+sT Xj+sT  1Ij   Xj+sT >↵ˆ(2) = s=0 s=0 s=0(2)>:  ˆj,MQC =N 1 PXj+sT  ↵(1) PXj+sT 1I(1)  ↵(2) PXj+sT 1I(2) jjjjs=0 s=0 s=0for j = 1,...,T. The consistency and asymptotic distribution of the CLS-estimators are established in the result given below.Theorem 3.1. The CLS-estimators are strongly consistent and asymptotically◆2 >✓◆✓P,N 1> (2) (2)s=0 s=0> N 1 N 1 N 1> >N 1(N Xj2+sT 1Ij   Xj+sT 1Ijnormal, i.e.,pˆd  1 1n(✓ ✓) !N(0,V WV ), (3.2)where V and W are square matrices of order 3T defined by blocks of 3 ⇥ 3 given by26 : 1 0 . . . 0 37 26 ⌦ : 1 0 . . . 0 376 0 :2 ... 0 7 6 0 ⌦:2 ... 0 7V=64 . . ... . 75 and W=64 . . ... . 75, (3.3)where0 0 ... :T 0 0 ... ⌦:T (k,l):j := E @ gj(✓j,Xj+sT 1) @ gj(✓j,Xj+sT 1) ,@✓k:j @✓l:j:=EU2 @ g(✓,X ) @ g(✓,X) ,k,l 2 {1,2,3}; j = 1,...,T and ✓j := (✓1:j,✓2:j,✓3:j) ⌘ ⇣↵(1),↵(2), j⌘, are⌦are the elements of the matrices :j and ⌦:j , j = 1, . . . , T , respectively, with(k,l):j j+sT @✓a:j j j j+sT  1 @✓l:j j j j+sT  1 jjthe parameters associated to j-th period.Proof. Consistency and asymptotic normality can be proved using the results of Klimko and Nelson (1978, sec.3). It is easily checked that all regularity con- ditions by Klimko and Nelson (1978, p. 634) are satisfied, and thus, by their]]></page><page Index="99" isMAC="true"><![CDATA[PERIODIC THRESHOLD PROCESSES 89Theorem 3.1 it follows that the CLS-estimators are strongly consistent. Fur- thermore, in proving asymptotic normality we have to check first conditions (A)-(C) in Monteiro et al. (2012, p. 2725). To this extent, note first that condi- tions (A) and (B) follow easily by the arguments given in Monteiro et al. (2012). Finally, note that each block :j of matrix V in (3.3) is defined aswhere:j=" q(1)m(1) j j,200 q(1)u(1) # j jq(2) m(2) q(2) u(2) , j j,2 j jq(1) u(1)jj jjq(2) u(2) 1u(1) :=E[Xj 1+sT|Xj 1+sT d rj] u(2) :=E[Xj 1+sT|Xj 1+sT d >rj]jhijhi m(1) := E Xi |Xj 1+sT d  rj m(2) := E Xi |Xj 1+sT d > rjj,i j  1+sT j,i j  1+sT q(1) := P [Xj+sT d  rj] q(1) := P [Xj+sT d > rj].jjNoticing that det| :j| =6 0, 8j = 1,...,T lead us to conclude that V is in- vertible and condition (C) is thus fulfilled. Thus, Theorem 3.2 of Klimko and Nelson (1978) is thereby satisfied implying that the CLS-estimators are asymp- totically normal. This concludes the proof. ⇤3.2. Conditional maximum likelihood estimators. For a fixed value of x0, the conditional likelihood function for the PSETINAR(2; 1, 1)T takes the formN 1 T QQs=0j=1 ⇣N 1 T QQs=0 j=1L(✓) := =Pj (Xj+sT = xj+sT |Xj 1+sT = xj 1+sT )⌘ ⇣⌘ x m(1) (1) (2) (2)pj xj 1+sT,xj+sT,↵j Ij +↵j Ij , jwith= pj (xj 1+sT , xj+sT )(1) + pj (xj 1+sT , xj+sT )(2) andM⇤ :=min(xj 1+sT,xj+sT).The CML-estimators✓ˆCML :=(↵ˆ(1) ,↵ˆ(2) , ˆ1,CML,...,↵ˆ(1) ,↵ˆ(2) , ˆT,CML), 1,CML 1,CML T,CML T,CMLpj ⇣xj 1+sT,xj+sT,↵(1)I(1) +↵(2)I(2), j⌘= jj jjM⇤ 2  j XX xj 1+sT (k)m (k) (xj 1+sT  m)   j+sT (k)Cm ↵j 1 ↵j (xj+sT  m)!Ij= e⌘ pj xj 1+sT , xj+sT , ↵(1)I(1),  j + pj xj 1+sT , xj+sT , ↵(2)I(2),  j⇣ m=0k=1 ⌘jj jj⇣ ⌘]]></page><page Index="100" isMAC="true"><![CDATA[90 I. PEREIRA, M. SCOTTO, AND R. NICOLETTEare obtained by maximizing the conditional log-likelihood function⇣⌘  ↵ˆ(1)   ↵(1)    11 PPN 1 T (1) (1) (2) (2)`(✓) =From the partial derivatives of first-order we obtain the set of systemss=0 j=18ln pj xj 1+sT,xj+sT,↵j Ij +↵j Ij , j .=0 ,> > > >N 1I(k) X (k)j (x x↵)  ↵(k)(1 ↵(k)) j+sT j 1+sT j> <p⇣j j s=0 ⌘ x ,x  1,↵(1)I(1) +↵(2)I(2), > >  jj j 1+sT j+sT j j j j j p ⇣x ,x ,↵(1)I(1) +↵(2)I(2),  ⌘>j j 1+sT j+sT j j j j j> >:P pjs=0 jjjjN 1 ⇣ (1) (1) (2) (2)⌘xj 1+sT ,xj+sT  1,↵j Ij+↵j Ij , j pj ⇣xj 1+sT ,xj+sT ,↵(1)I(1)+↵(2)I(2), j ⌘= Nfor k = 1,2 and j = 1,...,T. In order to solve those systems numerical proce- dures have to be employed. Note, however, that the CML-estimates for the  ’s, are readily available from those for the ↵’s through the following expression1 N 1 X⇣⌘s=0xj+sT  ↵ˆ(k) xj 1+sT , j=1,...,T. N j,CML ˆj,CML =The following result establishes the consistency and the asymptotic distributionof the CML-estimators.Theorem 3.2. Let (Xt ) be the PSETINAR(2; 1, 1)T model in (1.1). The CML-estimators are asymptotically normal, i.e,    ↵ˆ ( 2 )   ↵ ( 2 )     211 M0...03    ˆ1  1   6 1 7pn  .  !d N(0,I 1), where I=6 0 M2 ... 0 7 ,   .   4 . . ... . 5    ↵ˆ ( 1 )   ↵ ( 1 )     T T  00...MT  ˆT  T      ↵ˆ ( 2 )   ↵ ( 2 )     TT]]></page><page Index="101" isMAC="true"><![CDATA[PERIODIC THRESHOLD PROCESSES 91is the Fisher information matrix with26  E" @2`(✓) #  E" @2`(✓) #  E" @2`(✓) # 37 6 @(↵(1))2 @↵(1)@↵(2) @↵(1)@ j 7Mj = 6  E" @2`(✓) #  E" @2`(✓) #  E" @2`(✓) # 7,6 @↵(1)@↵(2) @(↵(2))2 @↵(2)@ j 7 6jj j j7jjjj64  E " @2`(✓) #  E " @2`(✓) # @ ↵ ( 1 ) @   j @ ↵ ( 2 ) @   jfor j = 1,...,T, and E @2`(✓)  75 @   2jh(2↵(k)  1)b a↵(k)2i j jE+1+1 (k) ⇢ @(↵(k))2 ↵(k)(1   ↵(k))jj"#@2`(✓) = NX XP (Xj+sT = b) Ijjabjj(k) (k)2 (k)⇥pj(b|a)+2(1 ↵j ) jpj(b 1|a) + jpj(b 2|a)  (k) (k) 2 (k) 2 p2j (b   1|a)(k)+2(1 ↵j ) jpj(b 1|a) + jpj(b 2|a)   j pj(b|a)(k) ; " # +1+1 ⇢  E @2`(✓) = N X XP(Xj+sT =b)  pj(b 1|a)(k)@↵(k)@ j ↵(k)(1   ↵(k)) jjjab(k) p2(b   1|a)(k)   j pj (b   2|a) +  j p(b|a)(k) ;+1+1 ( 2 (k) ) @   2j a b p j ( b | a ) ( k ) E @2`(✓) =NXXP(Xj+sT =b)  jpj(b 2|a)(k) pj(b 1|a) ,for k = 1, 2.Proof. In order to derive the large sample distribution of the CML-estimators, we use the same arguments as in Franke and Seligmann (1993, pp. 324–325). Note that the consistency and asymptotic distribution of the CML-estimators for the INAR(1) process can be obtained by means of Theorems 2.1 and 2.2 in Billingsley (1961, pp. 10–13). Since the innovation process is Poisson-distributed the arguments used by Monteiro et al. (2010, 2012) for the periodic INAR of order one and period T , PINAR(1)T , and the SETINAR(2; 1, 1) process can be easily generalized for the SETINAR(2; 1, 1) with periodic structure. We omit the details. ⇤]]></page><page Index="102" isMAC="true"><![CDATA[92 I. PEREIRA, M. SCOTTO, AND R. NICOLETTE4. Point Prediction in the PSETINAR modelIn this section we consider the forecasting of future values Xi+NT+h, with h = j + `T, ` 2 N0, given past observations up through time i + NT.First note that by iterating equation (1.1) it follows that Xt can be expressedas0n 1 YXYwhere, for t > in 1 0i 1 1 @AX@Aj=0 i=1 j=0 n 1⌘  t,n   Xt n +  t,i   Zt i, i=01dXt =  t j   Xt n +  t j  Zt i + Zt⇢ Qi 1  t j, i>0 ⇢  t,j T`,T, i=j+`T, j=1,...,T  t,i:= j=0 = ,1, i=0 1, i=0 leading to obtaind hX 1Xi+NT+h =  i+NT+h,h   Xi+NT +  i+NT+h,m   Zi+NT+h m. (4.1)m=0Since h = j + `T, it follows that`T 1 ` 1 T 1X  j+`T,m  Zj+`T m =d X X  j+`T,m T`,T  Zj+(` w)T m,m=0 w=0 m=0 the expression in (4.1) takes the formXi+NT+h =d  i+j+(N+`)T,j+`T  Xi+NT +Due to the periodicity of the  ’s,  i+j+(N+`)T,j+`T =  i+j,j+`T =  i+j,j T` ,T ,and considering the relationwithXi+NT+h =d   i+j,j T`,T  Xi+NT +Vi+j+`T i 1 ` 1 T 1(4.2)Vi+j+`T := X  i+j,m Zi+j m+NT +X X  i+j+(N+`)T,m+j+Tw Zi+(N+` w)T m. m=0 w=0 m=0j + X` T   1 m=0 i+j+(N+`)T,m  Zi+j+(N+`)T m.]]></page><page Index="103" isMAC="true"><![CDATA[PERIODIC THRESHOLD PROCESSES 93In order to generate the h-step ahead prediction the mean, median or mode of the predictive distribution of Xi+NT+h|Xi+NT can be employed as a point fore- cast. Note that the median and mode are considered as coherent (i.e., integer- valued) predictions, whereas the mean is not. In order to evaluate the prediction performance given by the mean, median or mode of the predictive distribution we can use the square root of the mean squared error (RMSE), the mean absolute error (MAE) or the loss function everything or nothing (LFEN), re- spectively. Note that the h-step-ahead point predictor that minimizes the mean square error (MSE) is given bywithXˆi+NT+h= E [Xi+NT+h|Xi+NT ] X⇥ `  ⇤= E  i+j,j T,T   Xi+NT |Xi+NT⇣⌘j 1+  i+j,m  (1) p(1) +  (2) p(2)min(x,y) 2 ⇣ ⌘ = P P Cx ↵(k)m=0 k=1 m j+1 j+1 j+NT+1m1   ↵(k) j+1x m e  (k) j+1(k) (y m)  j+1I(k) , j+1i+j  m i+j  m i+j  m i+j  m⇣⌘ (1) p(1) +  (2) p(2) i m i m i m i mp(1) := P (Xi+j m+NT 1  ri+j m); i+j  mm=0` 1 T 1+ X X  i+j,m+j+T w w=0 m=0p(2) = 1   p(1) , i + j > m;i+j m   i+j m p(1):=PXi m i+(N +` w)T  m 1p(2) =1 p(1) , i>m. i m i mr ; i mTurning now to the particular case h = 1, the one-step-ahead predictive func- tion is given byP(Xj+NT+1 = y|Xj+NT = x)⇣⌘⇣⌘(y m)!with  (k) = E[Z(k) ], k 2 {1,2}. Finally, from (4.2), the most commonlyused one-step-ahead predictor of Xj+NT+1, takes the form Xˆj+NT+1 = ⇣↵(1) Xj+NT + (1) ⌘P(Xi+NT ri+1)+ ⇣↵(2) Xj+NT + (2) ⌘P(Xi+NT >ri+1). j+1 j+1j+1 j+1]]></page><page Index="104" isMAC="true"><![CDATA[94 I. PEREIRA, M. SCOTTO, AND R. NICOLETTE5. Concluding RemarksThis paper has presented the periodic self-exciting threshold integer-valued autoregressive model of order one with period T , driven by a periodic sequence of independent Poisson-distributed random variables. Basic probabilistic and statistical properties of the model are established as well as parameter estima- tion and forecasting.We would like to stress here that an important issue when fitting PINAR models lies with parsimony. Even every simple PINAR model can have an in- ordinately large number of parameters. This is also true when dealing with PSETINAR models. Therefore, the development of procedures for dimensionality reduction is an impeding problem. This remains a topic of future research.References[1] P. Billingsley, Statistical Inference for Markov Processes, University of Chicago Press, Chicago, 1961.[2] J-P. Dion, G. Gauthier, and A. Latour, Branching processes with immigration and integer-valued time series, Serdica Math. J. 21, 123–136, 1995.[3] J. Franke and T. Seligmann, Conditional maximum likelihood estimates for INAR(1) processes and their application to modeling epileptic seizure counts, in: Developments in Time Series Analysis, T. Subba Rao (ed.), Chapman & Hall/CRC, London, 310–330, 1993.[4] J. Franke and T. Subba Rao, Multivariate first order integer valued autoregres- sions,Technical Report, Universitat Kaiserslautern, 1995.[5] L.A. Klimko and P. I.Nelson, On conditional least squares estimation for stochastic processes, Ann. Statist. 6, 629–642, 1978.[6] M. Monteiro, M. G. Scotto, and I. Pereira, Integer-valued autoregressive processes with periodic structure, J. Statist. Plann. Inference 140, 1529–1541, 2010.[7] M. Monteiro, M. G. Scotto, and I. Pereira, Integer-valued self-exciting threshold autore- gressive processes, Comm. Statist. Theory Methods 41, 2717–2737, 2012.[8] F. W. Steutel and K. van Harn, Discrete analogues of self-decomposability and stability, Ann. Probab. 7, 893–899, 1979.[9] K. F. Turkman, M. G. Scotto, and P. de Zea Bermudez, Non-Linear Time Series: Extreme Events and Integer Value Problems, Springer-Verlag, Switzerland, 2014.(I. Pereira and M. Scotto) Departamento de Matema´tica and CIDMA, Universidade de Aveiro, PortugalE-mail address: isabel.pereira@ua.pt; mscotto@ua.pt(R. Nicolette) CIDMAE-mail address: raquel.nicolette@gmail.com]]></page><page Index="105" isMAC="true"><![CDATA[A NOTE ON KERNEL ESTIMATION OF THE CONDITIONAL QUANTILE FUNCTION IN CONTINUOUS TIME ERGODIC PROCESSESANA CRISTINA ROSA AND MARIA EM´ILIA NOGUEIRADedicated to Nazar´e Mendes Lopes for everything she taught us and for her friendshipAbstract. We prove the strong consistency of a kernel type estimator of the conditional distribution function when the data generation process is time continuous and satisfies a general ergodic hypothesis. The strong consistency of the quantile function corresponding to this estimator is then established.1. IntroductionIt is well known that nonparametric methods provide essential tools for the statistical analysis and applications of stochastic processes, namely the pre- diction of future observations. Although most popular predictors are based on the regression function, other conditional parameters such as the mode and the quantiles have received considerable attention in literature. In particular, the robustness of quantiles to outliers and heavy-tailed error distributions jus- tifies their use as alternatives to the regression function for quantifying the dependence structure between a response variable and a covariate.Quantile estimates are usually derived either via quantile regression models (cf. Koenker [22], for an overview, Gannoun et al. [20] and Ghouch and Keilegom [21], among others) or by inverting nonparametric estimators of the conditional distribution function. Following this approach, the main results con- cern essentially the strong consistency and asymptotic normality of kernel type estimators and were obtained in the framework of time series under mixing as- sumptions on the data generation process (cf., for instance, Gannoun et al. [19] for complete data and Liang and Un˜a-A´lvarez [28] for censored observations).Accepted: 17 February 2015.2010 Mathematics Subject Classification. 62G05, 62M09, 62M20.Key words and phrases. nonparametric estimation, continuous time processes, ergodic data, kernel method, conditional quantile, strong consistency.The work was supported by the Department of Mathematics of the University of Coimbra.95]]></page><page Index="106" isMAC="true"><![CDATA[96 A. C. ROSA AND M. E. NOGUEIRAMore recently, the nonparametric estimation of conditional quantiles has been extended, in the same framework, to include functional covariates. For a general introduction to the subject, we refer to Ferraty and Vieu [18]. Several authors analysed the asymptotic properties of standard estimators, establi- shing almost complete consistency (Ferraty et al. [17]), normality (Ezzahrioui and Ould-Sa¨ıd [16]) and Lp norm consistency (Laksaci et al. [26], Dabo-Niang and Laksaci [9]).As far as we know, there are remarkably few results under a general depen- dence condition of ergodicity (cf. Delecroix et al. [11], Delecroix and Rosa [12], Yakowitz et al. [29], La¨ıb and Ould-Sa¨ıd [23]). However, after the work of La¨ıb and Louani [24], there has been a renewed interest in this subject, motivating some new consistency results with rate of convergence (cf. La¨ıb and Louani [25], Chaubey et al. [8], Chaouch and Khardani [7]).With regard to continuous time stationary processes, Banon [1], inspired by the discrete case, was the first author to propose a kernel density estimator based on an observed sample (Xt,0  t  T) of a di↵usion process X. Apart from di↵usion processes, we point out the pioneering work of Delecroix [10], who obtained rates of convergence for the MSE and the supremum norm cri- teria for a large class of density estimators, as well as the historical paper of Castellana and Leadbetter [6], addressing the convergence rate and the asymptotic normality of an estimator built by the delta method, in the set- ting of mixing processes. Since then, nonparametric estimation for continuous time processes has been extensively studied, proceeding roughly along the same line as for time series. Again, the asymptotic normality and improved rates of convergence of density and regression function estimators were obtained for several convergence criteria. Let us mention, in this context, the surveys of Bosq [4], Bosq and Blanke [5], and the references therein, as well as the works of Lejeune [27], focused on the histogram and the regressogram, and Blanke and Bosq [3], devoted to the regression function. Notice that the referred results were established by imposing mixing conditions on the underlying process, that are not satisfied in many cases, as pointed out by Didi and Louani [14]. We find in Didi and Louani [14, 15] interesting results concerning almost sure con- vergence of the kernel density and regression estimators, both pointwise and uniform, with nonparametric rates, under a continuous version of the ergodic hypothesis considered by Delecroix et al. [11]. In the same setting, Didi [13] proved the strong uniform consistency of a predictor based on the conditional mode.Following the work of the previous authors, we study, in the present work, the consistency of a kernel type estimator of the conditional distribution function.]]></page><page Index="107" isMAC="true"><![CDATA[ESTIMATION OF THE CONDITIONAL QUANTILE FUNCTION 97By inverting this estimator, we then derive a natural estimate of the conditional quantile function and its consistency is established.2. Assumptions and notationLet {(Xt, Yt) , t 2 [0, +1[ } be a Rd ⇥ R-valued continuous time measurable stochastic process on the probability space (⌦, A, P ).It is assumed that {(Xt, Yt) , t 2 [0, +1[ } is a strictly stationary and ergodic process with absolutely continuous marginal distributions.The density function of Xt will be denoted by g and, for each x 2 Rd such that g(x) > 0, fYt/Xt(·/x) ⌘ f(·/x) and FYt/Xt(·/x) ⌘ F(·/x) denote the density and the distribution function of Yt given Xt = x, respectively. We also considerq↵(x)=inf{z2R:F(z/x) ↵}, ↵2]0,1[,the conditional quantile function of Yt given Xt = x. Our estimator of FZ(·/x) is defined by+1 d FT(y/x)= 1] 1,y](z)fT(z/x)dz, y2R, x2R , 1where fT (·/x) is the kernel estimator of the conditional density fT (·/x) pro- posed by Didi [13]:1 Z T ✓ z   Yt ◆ ✓ x   Xt ◆ d+1 K0 K dtfT(z/x)=ThT0 hT hT, z2R.1 ZT ✓x Xt◆d K dtThT0 hTAs usual, K0 and K are density functions on R and Rd, respectively, and hT ⌘ h(T ), T > 0, is a real positive function.More precisely, FT takes the formFT(y/x)= NT (x,y),withand1 Z T ✓ y   Yt ◆ ✓ x   Xt ◆NT (x,y)=DT(x)=dK dt,DT (x)K dt ThT0 hT hTd G1 ZT ✓x Xt◆ThT0 hTwhere G is the distribution function corresponding to K0.]]></page><page Index="108" isMAC="true"><![CDATA[98 A. C. ROSA AND M. E. NOGUEIRAFrom the result obtained for FT (·/x), we deduce the consistency of the following estimator of q↵(x):q↵,T(x)=inf{z2R:FT(z/x) ↵}, ↵2]0,1[. Let us begin by introducing the following  -fields:• Gt = {Xs :0s<t},t>0,andG0 = (X0);• Ft = {(Xs,Ys):0s<t},t>0,andF0 = (X0,Y0);• St,  =  {(Xs,Ys),Xr :0s<t, trt+ }, t > 0,   > 0, and S0,  = {(X0,Y0),Xr :0r }, >0.In the sequel, E will denote the real interval [a, b], a < b.For easy reference, the assumptions needed to derive the announced results are gathered thereafter.(H1) Forevery t>0,  >0,(i) the conditional density of Xt with respect to Gt  , fGt  , existsXt and is lipschitzian on Rd (if t     < 0, take Gt   = G0);(ii) the conditional density of Xt with respect to Ft  , fFt  , existsXt(iii) the conditional density of (Xt,Yt) with respect to Ft  , fFt   ,(ift  <0,takeFt   =F0); exists; Z T(Xt ,Yt )(ift  <0,takeSt  ,  =S0, ).(H3) Forallx2Rd,thereexistrealconstantsC>0, >0and >0suchthat, for all (y1, y2) 2 E2 and (x1, x2) 2 V(x) ⇥ V(x), we have|F(y1/x1) F(y2/x2)|C |y1  y2|  +kx1  x2k  ,where V(x) denotes a neigbourhood of x and k · k the euclidean normof Rd.(H4) The kernel K verifies:(i) K is lipschitzian and has compact support, k⇤ = sup |K(x)|;(iv) lim 1 fGt  (x)dt=g(x)a.s., x2Rd. T!+1T0 Xt(H2) 8t 0, 8 >0, 8y2E,E✓G✓y Yt ◆.St  , ◆=E✓G✓y Yt ◆.Xt◆ a.s. hT hTZ(ii)kxk K(x) dx < +1.x2RdRd]]></page><page Index="109" isMAC="true"><![CDATA[ESTIMATION OF THE CONDITIONAL QUANTILE FUNCTION 99(H5) The kernel K0 verifies:(i) K0 is bounded, k0 = sup |K0(y)|;y2E (ii) K0 haZs compact support;+1  (iii) μ  = |u| K0(u) du < +1. 1T h2d(H6) lim hT =0, lim T =+1.T!+1 T!+1 logTNotice that hypothesis (H1) relies on the ergodic character of the data and was introduced by Didi and Louani [14, 15], as far as we know. The arguments presented by the authors show that it consists of a natural continuous-time version of the conditions considered by Delecroix et al. [11]. The interested reader is referred to Bergelson et al. [2] for some recent asymptotic results concerning continuous-time ergodic processes. Conditions (H4) to (H6) are quite common in kernel estimation of conditional functionals. The regularity assumption on the conditional distribution function imposed in (H3) is weaker than hypothesis (A3)(ii) of Chaouch and Khardani [7]. Assumption (H2) is a Markov-type condition similar to the one required in the previous paper (cf. (A4), p. 69) and to condition (H3) (i) of Didi and Louani [15].3. Main resultsNext we state the uniform almost sure convergence of FT (·/x) to the distri-bution function F(·/x).Theorem 3.1. Under the general conditions of the previous section and as-sumptions (H1) to (H6), we havesup |FT (y/x)   F (y/x)|  ! 0 a.s..y2E T !+1Proof. Let us consider:1 ZT ✓ ✓x Xt◆. ◆DT(x)=Thd EKhT Ft  dt; T01ZT ✓✓y Yt◆✓x Xt◆. ◆ NT(x,y)= d EG K Ft   dt;BT (x,y)= NT (x,y)  F(y/x). DT (x)T hT 0 hT hT]]></page><page Index="110" isMAC="true"><![CDATA[100A. C. ROSA AND M. E. NOGUEIRASinceFT(y/x) F(y/x) = BT (x,y)+ 1  NT (x,y) NT (x,y) ++ 1 DT (x)y2EDT (x)(BT (x,y)+F(y/x)) DT (x) DT (x) ,we may writesup|FT (y/x)   F (y/x)|1 A2,T (x) + DT (x)with A1,T(x) = sup|BT (x,y)|, A2,T(x) = sup NT (x,y) NT (x,y)  andA1,T (x) + + 1(A1,T (x) + 1) A3,T (x),   y2E   y2ETaking into account Lemma 4.1 and the fact that g(x) > 0, in order to prove Theorem 3.1 it is enough to establish the almost sure convergence to zero of Ai,T (x), for i = 1, 2, 3.(↵) We firstly analyse the behavior of A1,T (x).Note that A1,T (x) = 1 sup |CT (x, y)|, whereDT (x) y2ECT (x,y)=NT (x,y) F(y/x)DT (x).Equivalently, CT (x, y) is given by1 ZT ✓ ✓x Xt◆ ✓ ✓y Yt◆. ◆  . ◆d E K E G St  ,   F(y/x) Ft   dt. ThT 0 hT hTDT (x) A3,T (x) =  DT (x)   DT (x) .Then, (H2) implies that1ZT✓✓x X◆ ◆CT (x,y)= d E K t  (Xt)/Ft   dt, ThT 0 hTwith (Xt)=E✓G✓y Yt ◆/Xt◆ F(y/x). hTBut, for every w 2 ZRd such that g(w) > 0, we have +1 (w) = K0 (u) (F (y   hT u/w)   F (y/x)) du  1]]></page><page Index="111" isMAC="true"><![CDATA[ESTIMATION OF THE CONDITIONAL QUANTILE FUNCTIONand so, using (H3), |CT (x, y)| is almost surely bounded by101dt.CZT ✓✓x Xt◆⇣   ⌘  ◆ d E K hT μ  + kXt   xk Ft  ThT 0 hT Finally, under (H4), we get|CT (x,y)|DT (x)O⇣h T +h T⌘ a.s..The a.s. convergence to zero of A1,T (x) follows now from assumption (H6).( ) In order to prove the a.s. convergence to zero of A2,T (x), define, for   > 1,8>< b   a b   a> if 2N p = > h  T h  TT >b a  b a: h  +1if h  2/NTTwhere [u] denotes the integer part of the real number u. Dividing E = [a, b] into the intervalsIj,T =⇥a+(j 1)h T,a+jh T⇥, j2{1,...,pT  1}, andwe may writeA2,T (x)IpT,T =⇥a+(pT  1)h T,b⇤,max sup  NT (x, y)   N T (x, y) j2{1,...,pT } y2Ij,Tmax  NT (x,yj,T) NT (x,yj,T) ,j2{1,...,pT }max sup  NT (x,yj,T) NT (x,y) ,j2{1,...,pT } y2Ij,TwithA(1) (x)= 2,TA(2)(x)= 2,TA(3)(x)= 2,T2,T 2,T 2,Tmax sup |NT (x,y) NT (x,yj,T)|,=j2{1,...,pT } y2Ij,TA(1) (x) + A(2) (x) + A(3) (x),yj,T beinganarbitrarypointinIj,T,j2{1,...,pT}.The fact that G is absolutely continuous and (H5) (i) yield!(1)A2,T (x)  max supj2{1,...,pT } y2Ij,T= DT(x)O h  1 , T1d T hTZ T 0K✓ x   Xt ◆ | y   yj,T | hT hTk0 dt]]></page><page Index="112" isMAC="true"><![CDATA[102 A. C. ROSA AND M. E. NOGUEIRA as well as A(3) (x)  DT (x) O  h  1 .2,T TOn the other hand, recalling the notation and the result established inLemma 4.2, we conclude that, for every ✏ > 0,⇣ (2) ⌘P A2,Tn(x)>✏ =pTn max P j2{1,...,pTn}8(b a) h   exp Tn Z  !    Tn (1)  8(b a)h  T C1✏2 Tn nTn h2dTn .  Wt,Tn(x,yj,Tn)dt  > ✏   0   C1 ✏2 Tn h2d Tnlog TnThen, for T large enough, P ⇣A(2) (x) > ✏⌘ is bounded by C T    C1 ✏2L1 ,2 n2d under assumption (H6), for some positive constants C2 and L1.n2,TnA suitable choice of   and L1 assures that Tn 2d < +1, implying n=12,Tn( ) To conclude the a.s. convergence to zero of A3,Tn(x) it su ces to applyLemma.Lemma 4.2 with i = 0 and to use a similar reasoning. ⇤ We now state our main result.Theorem 3.2. Under the assumptions of Theorem 3.1, assume that f(·/x) is bounded. For all ↵ 2]0, 1[, we have|q↵,T (x)   q↵ (x)|  ! 0 a.s.. T !+1Proof. Observe that, for each ↵ 2 ]0, 1[, we may choose a, b 2 R such that q↵(x) 2 [a,b]. Using this fact and the boundness of f(·/x), we get|q↵,T (x)   q↵ (x)| = O( sup |FT (y/x)   F (y/x)|) a.s., y2[a,b]by the argument presented in [7, Lemma 6.7].4. Auxiliary results⇤Lemma 4.1. Suppose that conditions (H1) (i), (iv), (H4) (ii), (H5) (ii) hold. If K has compact support, we have|DT (x) g(x)|  ! 0 a.s. and  DT (x) g(x)   ! 0 a.s.. T !+1 T !+1Proof. Please see [13, p. 78]. ⇤P 2 +1  C1 ✏ L1the a.s. convergence of A(2) (x) to zero, as n ! +1, via the Borel-Cantelli]]></page><page Index="113" isMAC="true"><![CDATA[ESTIMATION OF THE CONDITIONAL QUANTILE FUNCTION 103Before stating the next lemma, we must fix some additional notations. For i = 0, 1, put ⇣ . ⌘W(i)(x,y)=Z(i)(x,y) E Z(i)(x,y) Ft   , t,T t,T t,Twhere ✓ ◆✓ ◆Z(i)(x,y)= 1 Gi y Yt K x Xt , y2E, x2Rd.t , T T h dT h T h T DenotingbyntheintegerpartofT,n=[T],T  1,let besuchthatn =T⌘Tn (andso1 <2).Wethendefine,fork2{0,...,n}, Tk =k  and Fk,  = {(Xs,Ys):0sTk}.Lemma 4.2. IfZK is a bounded kernel, we have, f or all n 2 N and ✏ > 0,  Tn(i)  ! 22dP   Wt,Tn(x,y)dt >✏ 4 exp  C1✏ Tn hTn 0 S with C1 a positive constant.nProof. Fix i in {0, 1}. Writing [0, Tn] = [Tk 1, Tk], we getZ T nW (i)k=1Xn (x, y) dt =XnU (i) k,TnV(i) (x,y)=Z Tl ⇣E⇣Z(i) (x,y).Fl 1, ⌘ E⇣Z(i) (x,y).Ft  ⌘⌘dt. l 1,Tn t,Tn t,Tn(x, y) +with ZTk⇣ ⇣.⌘⌘0 t,TnU(i) (x,y)= Z(i) (x,y) E Z(i) (x,y) Fk 1,  dt,k,Tn t,Tn t,Tn Tk 1Note that U(i)k,Tn k=1,...,n ` 1,Tn `=2,...,nrespectively, verifyingk=1l=2⇣Tl 1 ⌘ ⇣ ⌘V (i) l 1,Tnand V(i) (x,y) are martingale(x,y)di↵erences with respect to the   fields (Fk, )k=1,...,n and (F` 1, )`=2,...,n,(x, y)and U(i) (x,y)   2k⇤   1 , k , T n T n h dT n V (i) (x, y)   2 k⇤   1 , ` 1,Tn Tn hdTnk = 1,...,n, ` = 2, . . . , n.]]></page><page Index="114" isMAC="true"><![CDATA[104A. C. ROSA AND M. E. NOGUEIRAHence, Azuma’s inequality (cf. Delecroix et al. [11]) leads to    Xn     ✏ ! P   U(i) (x,y) >( ✏ 2 T n 2 h 2 d )  2exp   TnaswellasP=2exp  Tn , 32 k⇤2   Xn`=2  ✏! (x, y)  >(( ✏2Tnh2d))⇤   V (i) ` 1,Tn✏2 Tn2 h2d Tn  k,Tn   2 k=132k⇤2n 2( ✏2Tnh2d)2exp  Tn .2 exp 32 k⇤2 (n   1)  2 32 k⇤2  Then, taking C1 = 1 32 k⇤2    2, we have the announced result.5. Final remarksA straightforward improvement of Theorem 3.2 is the almost sure consis- tency of q↵,T(x), uniformly on x belonging to a compact subset of Rd. In line with the works of Didi and Louani [15] and Didi [13], we would deduce the almost sure convergence of the corresponding nonparametric predictor. In addition, it will be interesting to study consistency properties with rates of convergence.AcknowledgmentsWe kindly acknowledge the support of the Department of Mathematics of the University of Coimbra. We are also grateful to the referee for his valuable suggestions and comments which greatly improved the text.References[1] G. Banon, Nonparametric identification for di↵usion processes, SIAM J. Control Optim. 16 (3), 380–395, 1978.[2] V. Bergelson, A. Leibman, and C. G. Moreira, From discrete to continuous-time ergodic theorems, Ergodic Theory Dynam. Systems 32 (2), 383–426, 2012.[3] D. Blanke and D. Bosq, Regression estimation and prediction in continuous time, J. Japan Statist. Soc. 38 (1), 15–26, 2008.[4] D. Bosq, Nonparametric statistics for stochastic processes, Springer-Verlag, New York, 1998.[5] D. Bosq and D. Blanke, Inference and prediction in large dimensions, John Wiley and Sons, Chichester, 2007.]]></page><page Index="115" isMAC="true"><![CDATA[ESTIMATION OF THE CONDITIONAL QUANTILE FUNCTION 105[6] J. V. Castellana and M. R. Leadbetter, On smoothed probability density estimation for stationary processes, Stochastic Process. Appl. 21 (2), 179–193, 1986.[7] M. Chaouch and S. Khardani, Kernel-smoothed conditional quantiles of randomly cen- sored functional stationary ergodic data, J. Nonparametr. Stat. 27 (1), 65–87, 2015.[8] Y. Chaubey, N. La¨ıb, and A. Sen, Generalised kernel smoothing for nonnegative sta-tionary ergodic processes, J. Nonparametr. Stat. 22 (8), 973–997, 2012.[9] S. Dabo-Niang and A. Laksaci, Nonparametric quantile regression estimation for func-tional dependent data, Comm. Statist. Theory Methods 41 (7), 1254–1268, 2012.[10] M. Delecroix, Sur l’estimation des densit´es d’un processus stationnaire a` temps continu,Pub. Inst. Stat. Univ. Paris 25, 17–39, 1980.[11] M. Delecroix, M. E. Nogueira, and A. C. Rosa, Sur l’estimation de la densit´ed’observations ergodiques, Statist. Anal. Donn´ees 16 (3), 25–38, 1992.[12] M. Delecroix and A. C. Rosa, Ergodic processes prediction via estimation of the condi-tional distribution function, Ann. I.S.U.P. 39 (2), 35–56, 1995.[13] S. Didi, Quelques propri´et´es asymptotiques en estimation non param´etrique de fonc-tionnelles de processus stationnaires en temps continu, Ph.D. Thesis, University Pierreand Marie Curie - Paris VI, 2014.[14] S. Didi and D. Louani, Consistency results for the kernel density estimate on continuoustime stationary and dependent data, Statist. Probab. Lett. 83 (4), 1262–1270, 2013.[15] S. Didi and D. Louani, Asymptotic results for regression function estimate on continuoustime stationary and ergodic data, Stat. Risk Model. 31 (2), 129–150, 2014.[16] M. Ezzahrioui and E. Ould-Sa¨ıd, Asymptotic results of a nonparametric conditional quantile estimator for functional time series, Comm. Statist. Theory Methods 37 (17),2735–2759, 2008.[17] F. Ferraty, A. Rabhi, and P. Vieu, Conditional quantiles for functional dependent datawith application to the climatic El Nin˜o phenomenon, Sankhya¯ 67 (2), 378–398, 2005.[18] F. Ferraty and P. Vieu, Nonparametric functional data analysis, Springer-Verlag, Berlin,2006.[19] A. Gannoun, J. Saracco, and K. Yu, Nonparametric prediction by conditional medianand quantiles, J. Statist. Plann. Inference 117 (2), 207–223, 2003.[20] A. Gannoun, J. Saracco, A. Yuan, and G. E. Bonney, Non-parametric quantile regressionwith censored data, Scand. J. Statist. 32 (4), 527–550, 2005.[21] A.E.GouchandI.V.Keilegom,Locallinearquantileregressionwithdependentcensoreddata, Statist. Sinica 19, 1621–1640, 2009.[22] R. Koenker, Quantile regression, Cambridge University Press, Cambridge, 2005.[23] N. La¨ıb and E. Ould-Sa¨ıd, A robust nonparametric estimation of the autoregressionfunction under an ergodic hypothesis, Canad. J. Statist. 28 (4), 817–828, 2000.[24] N. La¨ıb and D. Louani, Nonparametric kernel regression estimation for functional sta- tionary ergodic data: asymptotic properties, J. Multivariate Anal. 101 (10), 2266–2281,2010.[25] N. La¨ıb and D. Louani, Rates of strong consistencies of the regression function estimatorfor functional stationary ergodic data, J. Statist. Plann. Inference 141 (1), 359–372,2011.[26] A. Laksaci, M. Lemdani, and E. Ould-Sa¨ıd, A generalized L1 approach for a kernelestimator of conditional quantile with functional regressors: consistency and asymptoticnormality, Statist. Probab. Lett. 79 (8), 1065–1073, 2009.[27] F. X. Lejeune, Histogramme, r´egressogramme et polygone de fr´equences en temps con-tinu, Ph.D. Thesis, University Pierre and Marie Curie - Paris VI, 2007.]]></page><page Index="116" isMAC="true"><![CDATA[106[28] [29]A. C. ROSA AND M. E. NOGUEIRAH. Liang and J. de Un˜a-A´ lvarez, Asymptotic properties of conditional quantile estimator for censored dependent observations, Ann. Inst. Statist. Math. 63 (2), 267–289, 2011. S. Yakowitz, L. Gyo¨rfi, J. Kie↵er, and G. Morvai, Strongly-consistent nonparametric forecasting and regression for stationary ergodic sequences, J. Multivariate Anal. 71 (1), 24–41, 1999.(A. C. Rosa and M. E. Nogueira) Departamento de Matema´tica, Apartado 3008, EC Santa Cruz, 3001   501 Coimbra.E-mail address: cristina@mat.uc.pt; memn@mat.uc.pt]]></page><page Index="117" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACHMARIA EDUARDA SILVANot everything that counts can be counted, and not everything that can be counted counts Thank you, Nazar´eAbstract. Timeseriesofcountsarisewhentheinterestliesonthenum- ber of certain events occurring during a specified time interval. Many of these data sets are characterized by low counts, asymmetric distributions, excess zeros, over dispersion, ruling out normal approximations. Several approaches and diversified models that explicitly account for the dis- creteness of the data have been considered in the literature, among which are the INteger-valued AutoRegressive, INAR, models. These models are based on random operations which, operating on discrete variables ensure an integer-valued result. The INAR models are attractive since they are linear-like models for discrete time series and exhibit recognizable corre- lation structures. This paper considers INAR models for analyzing time series of counts and discusses associated statistical inference, comprising estimation, diagnostics and model assessment.1. IntroductionThe problem of modelling time series of low counts has attracted many researchers over the last few decades. In fact time series of counts arise in many di↵erent contexts, usually as counts of certain events or objects in specified time intervals as for example: social science [31, 37], queueing systems [1], experimental biology [54], environmental processes [48, 11, 24], economics and finance [6, 39, 15, 28, 27], epidemiology [10, 53], international tourism demand [8, 17, 7], statistical control processes [50, 52], telecommunications [30], optimal alarm systems [34], and in the biopharmaceutical industry [3].Accepted: 23 February 2015.2010 Mathematics Subject Classification. 60G10, 62M10. Key words and phrases. Time Series, Count Data, INAR.The work was supported by Portuguese funds through the CIDMA - Center for Re- search and Development in Mathematics and Applications, and the Portuguese Foundation for Science and Technology (FCT Fundac¸˜ao para a Ciˆencia e a Tecnologia), within project UID/MAT/04106/2013.107]]></page><page Index="118" isMAC="true"><![CDATA[108 M. E. SILVAIn many cases, the discrete variates are large numbers and it may make sense to approximate them by continuous variates. Often, however, this is not possible or even desirable and it is necessary to develop an appropriate mod- elling strategy for the statistical analysis of time series of counts. One of the approaches developed is based on a random operation called thinning opera- tion capable of preserving the integer valued nature of the variables, giving rise to the class of INteger valued AutoRegressive, INAR, models. This paper considers the first order INAR, INAR(1), models for analyzing time series of counts and discusses associated statistical inference, comprising estimation, di- agnostics and model assessment. The plan of the paper is as follows. Section 2 introduces the INAR(1) models with including discussion of the most relevant properties. Section 3 considers parameter estimation and presents a set of tools appropriate to check the adequacy of fitted INAR models, an important part of any iterative modelling exercise in applied time series analysis. Section 4 illustrates the fitting of the models to a time series of counts of the stock type. Section 5 provides some concluding remarks.2. First order INteger valued AutoRegressive models Definition 2.1. The first order integer autoregressive, INAR(1), model is de-fined on the discrete support N0 by the recursive equationXt = ↵ ⇧ Xt 1 + ✏t, (2.1)where {✏t} is a sequence of independent and identically distributed non-negative integer valued random variables, for each t independent of Xt 1 and of ↵⇧Xt 1, with finite mean μ✏ and variance  ✏2 and, conditional on Xt 1, ↵ ⇧ Xt 1 is an integer valued random variable whose probability distribution depends on the parameter ↵.1Thus, 0⇧02 denotes a random operator, usually called thinning operator, which always produces integer values and introduces serial dependence via the conditioning on Xt 1. Consider now that in model (2.1) we require that the marginal distribution of {Xt}t is of the same family as {✏t}. [25] proposes an approach to solve this problem within the convolution-closed infinitely divisible class of (marginal) distributions. The random operator 0⇧0 is required not only to introduce serial dependence and preserve the integer-valued status of the random variable but also to be unconditionally of the same family as {✏t}t.1In fact, ↵ may be a vector of parameters but at this point we prefer the simpler scalar notation2The operator is in fact ↵⇧ as it depends on ↵ but usually the simpler notation is used.]]></page><page Index="119" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH 109The intuition behind the operator 0⇧0, as described by that author is the fol- lowing.Let F✓ denote a convolution-closed infinitely divisible parametric family such that F✓1 ⇤F✓2 = F✓1+✓2 , with 0⇤0 denoting the convolution operator. Let Y1, Y2 be independent random variables each with distribution F(1 ↵)✓ (pmf f(1 ↵)✓(·)) and Y12 be a another random variable independent from Y1, Y2 with distribution F↵✓ . The distribution of Y12 given Y12 + Y1 = y is denoted by G↵✓,(1 ↵)✓,y and its pmf by g(·|y). [25] writes the joint distribution of (Xt,Xt 1) as being the sameasthatof(Y12+Y2,Y12+Y1),inwhichcaseY12 representsacommonlatent or unobserved component of the pair (Xt, Xt 1) that carries the dependence of the observations between two consecutive time periods and Yi, i = 1, 2 represent the arrivals in the model.[25] shows that the processes defined in (2.1) are Markov order 1, time reversible and stationary with non-negative serial dependence, ⇢k = ↵k. The transition probabilities are given byP(Xt = k|Xt 1 = `) = =miXn{k,`} j=0miXn{k,`} j=0g(j|`)P(✏t = k   j) g(j|`)f(1 ↵)✓(k   j).(2.2)Moreover, the INAR(1) model is a member of the class of conditional linear first order autoregressive, CLAR(1), models introduced by [20].2.1. The Poisson INAR(1). Consider the case where F✓ is Poisson ✓ =  /(1   ↵). Setting G↵✓,(1 ↵)✓,x as Binomial(x, ↵), leads to the most common thinning operator which is the binomial thinning, denoted by 0 0 and originally introduced by [46] to extend the notions of self-decomposability (DSD) and stability to integer-valued time series.Definition 2.2. Let X be a non-negative integer-valued random variable. Then, for any ↵ 2 [0, 1] define the binomial thinning operator asX i=1where {Yi}i is a sequence of independent and identically distributed Bernoulli random variables with P(Yi = 1) = ↵, called the counting series of ↵ X, which is also independent of X.For properties of the binomial thinning operation see [47, 51, 43, 44].↵   X :=Yi, (2.3)]]></page><page Index="120" isMAC="true"><![CDATA[110 M. E. SILVAThe binomial thinning based INAR(1) model was originally proposed by [2] and [32]. The conditional distribution of Xt given Xt 1, fXt|Xt 1(xt|xt 1) = p(Xt|Xt 1) is now the convolution of the two components, binomial and Pois- son, as follows:i X  i e    Xt i↵ (1 ↵) t 1 , (2.4)= exp ⇢    [2 ↵ (1 ↵)(s1 +s2) ↵s1s2]  (2.5) 1 ↵is symmetric in its arguments s1 and s2 and therefore the process is time- reversible. The PoINAR(1) may be interpreted as an infinite server queue. The service time is geometric with parameter 1 ↵ and the arrival process is Poisson with mean  . A fundamental result in queueing theory, Little Flow’s equation, states that the expected length of the queue is equal to the arrival rate times the expected waiting time. It is thus possible to compute the expected number of time units a newly arrival stays in the system and which is given by: 1/(1   ↵).3. Parameter Estimation and Diagnostic Tools3.1. Parameter estimation. This section considers the estimation of the pa- rameters in the INAR(1) models discussed previously. Estimation can be car- ried out in several ways leading to the following broad categories of estimators: moment based estimators (MM), regression based or conditional least squares (CLS) estimators and likelihood based (ML) estimators. All these approaches have been considered in detail in the literature for the Poisson model. Addition- ally Bayesian methodology has been considered by [36] and [45]. However, the most common approach for the estimation of the INAR(1) model is maximum likelihood method.Let x = (x1, . . . , xn) represent the observed time series and ✓ the s⇥1 vector of model parameters to be estimated. Since the INAR(1) model (2.1) is a first order stationary Markov chain, the likelihood function is written asYn t=2Mt✓ ◆ X Xt 1p(Xt|Xt 1)=where Mt = min (Xt 1, Xt). The bivariate pgf of the PoINAR(1) process, giveni=0 i(Xt   i)!by [4]PXt,Xt 1 (s1, s2) =PXt 1 (s1(1   ↵   ↵s2))P✏t (s2)Ln(✓|x) = P(X1 = x1)fXt|Xt 1 (xt|xt 1), (3.1)]]></page><page Index="121" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH 111where fXt|Xt 1(xt|xt 1) is given by (2.2) and P(X1 = x1) represents the sta- tionary marginal distribution. For asymptotic analysis the standard approach is to consider the conditional log-likelihood function given byXn   `n(✓)= log fXt|Xt 1(xt|xt 1) . (3.2)t=2Based on the results of [5] for Markov processes, [25, 26] prove the followingresultTheorem 3.1. The maximum likelihood estimators, MLE, of the parameters ✓ in INAR(1) models are consistent, asymptotically normal and asymptotically e cient.The proof requires that some regularity conditions hold ([26], pp. 318).1/2 ˆ  1 The limit distribution of n (✓   ✓) is N(0, ⌃M LE ), ⌃M LE = ⌃(✓) and⌃(✓) = ( ij (✓)) is a non-singular s ⇥ s matrix with elements   (✓) = E ⇣`˙ (✓;x ,x )`˙ (✓;x ,x )⌘.ij ✓i12j12In general, for numerical maximum likelihood estimation, a quasi-Newton method can be used with an input of the negative log-likelihood function and output of the MLE and inverse Hessian matrix at the MLE. The initial es- timates required by the optimization algorithm are based on the method ofˆmoments. The inverse Hessian evaluated at the maximum, ✓, can be used asˆ the estimated variance-covariance matrix of the ML estimator ✓.3.2. Diagnostic tools. A crucial step in any statistical investigation is the assessment of the adequacy of the models proposed and fitted to the data under analysis. Various methods for model validation and diagnostics in discrete- valued time series have been proposed in the literature. These methods can be broadly classified as: parametric resampling methods; residual based methods; methods based on the predictive distributions; model comparisons using scores and information criteria.3.2.1. Parametric resampling methods. [49] proposes a procedure based on para- metric bootstrap and special functionals designed to show the specific features of interest. Specifically, the fitted model is used to generate many samples, all with the same number of observations as the original data set. The samples generated are then used to construct an empirical distribution of the functional of interest. If the fitted model is adequate in describing the feature of inter- est, the functional quantity of the original data should be a reasonable point with respect to the empirical distribution. The functional of interest may be]]></page><page Index="122" isMAC="true"><![CDATA[112 M. E. SILVAa spectral density function or the autocorrelation function. Here, the autocor- relation properties are of interest. Thus M artificial data sets, with the same length of the original data set, are generated from the fitted model. Based on these, M sample autocorrelation functions (ACF) are obtained. For each fixed lag of the ACF, the (1   ↵/2) and ↵/2 quantiles of the empirical distribution, denominated acceptance bounds, are computed. Then, a probability interval is obtained for each lag and an envelope is obtained for the ACF. If the fitted model is adequate at each lag, the sample ACF of the original data should be largely within the envelope. Plotting the envelope and the sample ACF of the data jointly, gives rise to a graphical display that can be used to assess the overall goodness of fit of the fitted model with respect to the serial correlation properties. A model is considered to be adequately reproducing the correlation structure of the data, if the sample ACF of the observed data lies within these acceptance bounds. Note that since the sample ACF at di↵erent lags are corre- lated, the acceptance envelope considered is not a joint 100(1   ↵)% confidence interval of the sample ACF.3.2.2. Residual based methods. The dynamic structure in the mean and dis- persion properties may be checked using tools based on the Pearson residuals defined byrt = Xt   E(Xt|Xt 1), (3.3) Var(Xt |Xt 1 )1/2where the population quantities are replaced by their estimated counterparts. If the model is correctly specified, these residuals should exhibit mean zero and variance one and no (significant) serial correlation.However, the structure of the INAR(1) model suggests additional residu- als checks. In fact, the INAR(1) model maybe seen as a structural model in the sense that it considers the data to be composed of a set of unobserved components each of which captures a feature of the data: the first component ↵⇧Xt 1 specifies the random number departures or its complement the random number of survivors from the past while ✏t represents the new arrivals at the system at time t. This interpretation leads to a residual decomposition that allows to check the adequacy of each component. For details on the residual decomposition and subsequent testing procedures see [16].3.2.3. Methods based on the predictive distributions. A useful tool to check the adequacy of the distributional assumptions of the models is a suitably speci- fied and modified version of the probability integral transform, PIT, originally proposed by [40]. This device has been used in the assessment of predictive dis- tributions of a continuous type by [40], [14] and recently by [18] and [19]. The PIT has a uniform distribution when the underlying model is continuous. For]]></page><page Index="123" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH 113discrete valued variables the associate distribution functions are step functions, therefore adjustments are necessary. Several authors propose a randomized PIT obtained by perturbing the step function nature of the distribution function of the discrete random variables. Additionally, [12] introduce a nonrandomized version of the PIT suitable to count data. For details and references see [12].Further evaluation of the model based on its predictive performance may be carried out using scoring rules suggested by [12] and [28].3.2.4. Information criteria. Akaike information criterion, AIC and its many variants has been one of the most popular tools for model selection in time se- ries analysis. In the context of time series of counts, [41] studies an automatic criterion for selecting the order of an INAR(p) model based on the corrected version of Akaike Information Criterion, AICC of [22]. Some authors have used AIC as means of choosing between non nested models for time series of counts, regardless of the lack of studies concerning the performance of the criterion in this framework. Moreover, [38] examine the ability of widely used information criteria such as AIC, BIC and the Hannan-Quinn criterion (HQ) [21] to dis- tinguish between some nonlinear times series models that have been popular with practitioners. After performing an extensive simulation study they argue that all three criteria have a useful role to play in a time series model selection exercise.4. IllustrationThis section illustrates the modelling procedure with a data set consisting of the number of di↵erent IP addresses accessing the server of the pages of the Department of Statistics of the University of Wu¨rzburg in two-minute periods from 10 am to 6 pm on the 29th November 2005, in a total of 241 observations. This data set was originally studied by [50] and exhibits small but significant autocorrelation as indicated by Figure 1. The sample mean and variance x = 1.31 and  ˆ2 = 1.39 do not indicate overdispersion.Fitting a PoINAR(1) model to the data yields the CML estimates ↵ˆ = 0.24(0.00) and  ˆ = 1.01(0.01). The parametric bootstrap exercise with M = 1000 and the residual analysis represented in Figure 2 indicate that the model captures the dynamics of the data. In fact, the variance of Pearson residuals is 1.05 and the acf of the component residuals in (c) indicate that the residuals are white noise. However, Figure 2(b) indicates that at time t = 224 the residual is unusually large with a large arrival component. This may suggest the occur- rence of an additive outlier meaning that Xt=224 may be contaminated by an exogenous source but the e↵ect is not carried over to subsequent observations by the dynamics.]]></page><page Index="124" isMAC="true"><![CDATA[114M. E. SILVA(a)0 50 100 150 200 250 t(b)5 10 15 20 lagFigure 1. Number of di↵erent IP addresses accessing the server of the pages of the Department of Statistics of the University of Wu¨rzburg between 10 am and 6 pm on 29 November 2005 (a) and corresponding sample autocorrelation (b).0.0 0.4 0.8 0 2 4 6 8]]></page><page Index="125" isMAC="true"><![CDATA[ MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH115     (a)(c)ResidualsACF−2 0 2 4 6−0.2−0.1 0.0 0.10.20.30.4 0.5−0.2−0.10.0 0.10.2 0.3Correlation                5 10 15 20lag510lagResidual Arrival Departure15 DataAcceptance boundsResiduals Arrival Departure    0 50 100 150 200tFigure 2. Parametric bootstrap exercise (a), component residu- als (b) and corresponding autocorrelations (c) for IP data set.[42] propose a Bayesian approach to model such outliers assuming that the observed process Yt is obtained from the unobservable clean process Xt con- taminating each Xt with probability  t with an outlier of random size ⌘t. ThusYt = Xt + ⌘t t,with Xt =↵⇧Xt 1 +et and  t ⇠Be(pt), (4.1)(b)]]></page><page Index="126" isMAC="true"><![CDATA[116 M. E. SILVAwhere  1, ⌘1, . . . ,  n, ⌘n are independent and independent of the latent process Xt and ⌘t, the random size of the outlier at time t is a random variable with the same support as Xt and mean   : ⌘t ⇠ Po( ). The Bayesian approach to estimate model (4.1) requires apriori distribution for the parameters of interest. For the parameters 0 < ↵ < 1 and   > 0 the traditionally weakly informative priors for the PoINAR(1) ([45]) are chosen: a non-informative Beta prior with parameters a = 0.01,b = 0.01 and a a non-informative Gamma prior with parameters c = 0.01, d = 0.01, respectively. The prior specifications for pt, the probability of contamination is a Beta distribution with parameters (g = 5, h = 95), with expectation E(pt) = 0.05, reflecting the belief that outliers occur occasionally. The prior for the mean size of the outliers   is a non informative Gamma distribution. The result of applying the outlier detection methodology is represented in Figure 3 indicating the occurrence of an outlier at time t = 224 with high probability.(a)0 50 100 150 200 250t(b)0 50 100 150 200tFigure 3. IP time series (a) and posterior probability of out- lier occurrence at each time (b).0.0 0.4 0.8 02468]]></page><page Index="127" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH 117The parameter estimates for model (4.1) are ↵ˆBayes = 0.27 and  ˆBayes = 0.89, with posterior distributions represented in Figure 4 and ⌘ˆ = 7 for the size of the outlier, leading to the following model:Xt = 0.27   Xt 1 + et, (a)et ⇠ P o(0.89).(4.2)Yt = Xt + 7I224,(b)0.20 0.25 0.30 0.350.70 0.800.90 1.00αλFigure 4. Posterior distribution of ↵ and  . The dotted lines represent the estimates ↵ˆBayes = 0.27 and  ˆBayes = 0.89.Figure 5 represents the residuals resulting from the fit of (4.2). Note that the largest residual reduces from 6.8 to 3.3, indicating a better fit.PAnfurther indication of the better fit is based on the prediction sum of squares t=2(yt   yˆt)2, where yˆt = E(yt|yt 1 = yt 1; parameter estimates) which drops from 317.9 to 264.0 when the outlier is included in the model.0 5 10 15DensityDensity0 2 4 6 8 10 12]]></page><page Index="128" isMAC="true"><![CDATA[118M. E. SILVA●● Residual ArrivalDeparture●● ●●● ●● ●●●●●●●●● ●● ● ●● ●●● ● ●●●●● ●●●● ●● ●●●●●●● ●●●●●       ●●●● ●● ● ● ●●● ● ●●  ●● ●●●● ●●●●●● ●●●●●● ● ●●● ● ●●●●● ●● ●●●      ● ●●●●●●● ●●● ●●●●● ● ● ● ● ●      ● ● ● ● ●       ●   ● ● ● ●●● ●●●● ● ●●● ●●●●●●●●●●●●●●● ●●●●● ●●● ●●● ●● ● ● ●● ●●●●●● ●●●●●●● ●●●●● ● ●●● ●●●●●●●●● ●●●●● ● ●●●●●●●0 50 100 150 200 tFigure 5. IP data set: residuals of a PoINAR(1) with an outlier at time t = 224.5. Final remarksTime series of counts arise in a wide variety of fields. The need to analyse such data adequately led to a multiplicity of approaches and a diversification of models that explicitly account for the discreteness of the data. One approach is based on the generalized linear models theory for dependent data [29]. Another point of view into the problem is given by parameter driven models which postulate that the observed process is driven by an unobserved process [13]. Yet another approach to the problem of modelling dependent count data is based on the use of renewal processes for generation of a correlated sequenceResiduals−1 0 1 2 3]]></page><page Index="129" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH 119of Bernoulli trials [11]. Here we focused on the INAR(1) models which are a class of observation-driven models particularly suited for stock type data. We illustrated the modelling of a time series with INAR(1) models.Several generalizations of the INAR(1) models are available in the literature, namely: INAR(p) models [23, 9], models with moving average components, INARMA, [33], periodic models [35] and bivariate models [37].References[1] S. Ahn, G. Lee, and J. Jeon, Analysis of the m/d/1-type queue based on an integer- valued first-order autoregressive process, Oper. Res. Lett. 27, 235–241, 2000.[2] M. Al-Osh and A. Alzaid, First-order integer-valued autoregressive (INAR(1)) process, J. Time Series Anal. 8, 261–275, 1987.[3] M. Alosh, The impact of missing data in a generalized integer-valued autoregression model for count data, J. Biopharm. Statist. 19, 1039–1054, 2009.[4] A. Alzaid and M. Al-Osh, First-order integer-valued autoregressive (INAR(1)) process: distributional and regression properties, Stat. Neerl. 42, 53–61, 1988.[5] P. Billingsley, Statistical Inference for Markov Processes, University of Chicago Press, 1961.[6] R. Blundell, R. Gri th, and F. Windmeijer, Individual e↵ects and dynamics in count data models, J. Econometrics 108, 113–131, 2002.[7] K. Bra¨nn¨as, J. Hellstr¨om, and J. Nordstro¨m, A new approach to modelling and fore- casting monthly guest nights in hotels, Int. J. Forecasting 18, 19 – 30, 2002.[8]K.Bra¨nn¨asand J.Nordstro¨m,Touristaccommodatione↵ectsoffestivals,Tourism Econ. 12, 291–302, 2006.[9] B. Bu and B. McCabe, Maximum likelihood estimation of higher-order integer-valued autoregressive process, J. Time Series Anal. 24, 973–994, 2008.[10] M. Cardinal, R. Roy, and J. Lambert, On the application of integer-valued time series models for the analysis of disease incidence, Stat. Med. 18, 2015–2039, 1999.[11] Y. Cui and R. Lund, A new look at time series of counts, Biometrika 96, 781–792, 2009.[12] C. Czado, T. Gneiting, and L. Held, Predictive model assessment for count data, Bio-metrics 65, 1254–1261, 2009.[13] R. Davis and R. Wu, A negative binomial model for time series of counts, Biometrika96, 735–749, 2009.[14] A. Dawid, Statistical theory: The prequential approach, J. Roy. Statist. Soc. Ser. A 147,278–292, 1984.[15] K. Fokianos, A. Rahbek, and D. Tjostheim, Poisson autoregression, J. Amer. Statist.Assoc. 104, 1430–1439, 2009.[16] R. K. Freeland and B. P. M. McCabe, Analysis of low count time series data by Poissonautoregression, J. Time Series Anal. 25, 701–722, 2004.[17] A. Garc´ıa-Ferrer and R. A. Queralt, A note on forecasting international tourism demandin Spain, Int. J. Forecasting 13 (4), 539–549, 1997.[18] T. Gneiting, F. Balabdaoui, and A. E. Raftery, Probabilistic forecasts, calibration andsharpness, J. R. Stat. Soc. Ser. B Stat. Methodol. 69, 243–268, 2007.[19] T. Gneiting and A. E. Raftery, Strictly proper scoring rules, prediction, and estimation,J. Amer. Statist. Assoc. 102, 359–378, 2007.]]></page><page Index="130" isMAC="true"><![CDATA[120 M. E. SILVA[20] G. Grunwald, R. Hyndman, L. Tedesco, and R.Tweedie, Non-Gaussian conditional linear AR(1) models, Aust. N. Z. J. Stat. 42, 479–495, 2000.[21] E. J. Hannan and B. G. Quinn, The determination of the order of an autoregression, J. R. Stat. Soc. Ser. B. Methodol. 41, 1979.[22] C. M. Hurvich and C. L. Tsai, Regression and time series model selection in small samples, Biometrika 76, 297–307, 1989.[23] D. Jin-Guan and L. Yuan, The integer-valued autoregressive INAR(p) model, J. Time Series Anal. 12, 129–142, 1991.[24] M. Scotto, C. Weiß, M. E. Silva, and I. Pereira, Bivariate binomial autoregressive models, J. Multivariate Anal. 125, 233 – 251, 2014.[25] H. Joe, Time series model with univariate margins in the convolution-closed infinitely divisible class, J. Appl. Probab. 33 (3), 664–677, 1996.[26] H. Joe, Multivariate Models and Dependence Concepts, Chapman & Hall//CRC, Lon- don, 1997.[27] R. C. Jung and A. R.Tremayne, Useful models for time series of counts or simply wrong ones?, AStA Adv. Stat. Anal. 95, 59–91, 2011.[28] R. C. Jung and A. R. Tremayne, Convolution-closed models for count time series with applications, J. Time Series Anal. 32, 268–280, 2011.[29] B. Kedem and K. Fokianos, Regression Models for Time Series Analysis, Wiley, Hobo- ken, NJ, 2002.[30] D. Lambert and C. Liu, Adaptive thresholds: monitoring streams of network counts, J. Amer. Statist. Assoc. 101, 78–88, 2006.[31] B. McCabe and G. Martin, Bayesian predictions of low count time series, Int. J. Fore- casting 21, 315 – 330, 2005.[32] E. McKenzie, Some ARMA models for dependent sequences of Poisson counts, Adv. in Appl. Probab. 20 (4), 822–835, 1988.[33] E. McKenzie, Discrete variate time series, in: Stochastic processes: modelling and sim- ulation, C. Rao, D. Shanbhag (eds.), Handbook of Statistics 21, Elsevier Science, Ams- terdam, 573–606, 2003.[34] M. Monteiro, I. Pereira, and M. G. Scotto, Optimal alarm systems for count processes, Commun. Stat. - Theor. M. 37, 3054–3076, 2008.[35] M. Monteiro, I. Pereira, and M. G. Scotto, Integer-valued autoregressive processes with periodic structure, J. Statist. Plann. Inference 140, 1529–1541, 2010.[36] N. Silva, Ana´lise bayesiana de s´eries temporais de valor inteiro, Tese de Doutoramento em Matem´atica, Universidade de Aveiro, Portugal, 2005.[37] X. Pedeli and D. Karlis, A bivariate INAR(1) process with application, Stat. Model. 11, 325–349, 2011.[38] Z. Psaradakis, M. Sola, F. Spagnolo, and N. Spagnolo, Selecting nonlinear time series models using information criteria, J. Time Series Anal. 30, 369–394, 2009.[39] A. Quoreshi, Bivariate time series modeling of financial count data, Comm. Statist. Theory Methods 35, 1343–1358, 2006.[40] M. Rosenblatt, Remarks on a multivariate transformation, Ann. Math. Statist. 23, 470– 472, 1952.[41] I. Silva, Analysis of discrete-valued time series: Some contributions to discrete-valued time series, LAP LAMBERT Academic Publishing, 2012.[42] M. E. Silva and I. Pereira, Detection of additive outliers in Poisson INAR(1) time series, in: Mathematics of Planet Earth: energy and climate change, International Conference]]></page><page Index="131" isMAC="true"><![CDATA[MODELLING TIME SERIES OF COUNTS: AN INAR APPROACH 121and Advanced School Planet Earth, Portugal, March 21-28, 2013, J. P. Bourguignon,R. Jeltsch, A. A. Pinto, M. Viana (eds.), 419–430, 2014.[43] M. E. Silva and V. Oliveira, Di↵erence equations for the higher-order moments andcumulants of the INAR(1) model, J. Time Series Anal. 25, 317–333, 2004.[44] M. E. Silva and V. Oliveira, Di↵erence equations for the higher-order moments andcumulants of the INAR(p) model, J. Time Series Anal. 26, 17–36, 2005.[45] I. Silva, M. E. Silva, I. Pereira, and N. Silva, Replicated INAR(1) Processes, Methodol.Comput. Appl. Probab. 7, 517–542, 2005.[46] F. Steutel and K. Van Harn, Discrete analogues of self-decomposability and stability,Ann. Probab. 7, 893–899, 1979.[47] K. F. Turkman, M. G. Scotto, and P. de Zea Bermudez, Non-Linear Time Series:Extreme Events and Integer Value Problems, Springer-Verlag, Switzerland, 2014.[48] P. Thyregod, N. Carstensen, H. Madsen, and K. Arnbjerg-Nielsen, Integer-valued autore- gressive models for tipping bucket rainfall measurements, Environmetrics 10, 395–411,1999.[49] R. S.Tsay, Model checking via parametric bootstraps in time series analysis, J. R. Stat.Soc. Ser. C. Appl. Stat. 41, 1–15, 1992.[50] C. H. Weiß, Controlling correlated processes of Poisson counts, Qual. Reliab. Eng. Int.23, 741–754, 2007.[51] C. H. Weiß, Thinning operations for modelling time series of counts: a survey, AStAAdv. Stat. Anal. 92, 319–341, 2008.[52] N. Ye, J. Giordano, and J. Feldman, A process control approach to cyber attack detec-tion, Commun. ACM 44, 76–82, 2001.[53] X. Yu, M. Baron, and P. K. Choudhary, Change-point detection in binomial thinningprocesses, with applications in epidemiology, Sequential Anal. 32, 350–367, 2013.[54] J. Zhou and I. Basawa, Least-squares estimation for bifurcating autoregressive processes,Statist. Probab. Lett. 74, 77 – 88, 2005.(M. E. Silva) Faculdade de Economia, Universidade do Porto & CIDMA E-mail address: mesilva@fep.up.pt]]></page><page Index="132" isMAC="true"><![CDATA[]]></page><page Index="133" isMAC="true"><![CDATA[ON THE EXTREMES OF STATIONARY GAUSSIAN RANDOM FIELDS UNDER STRONG DEPENDENCEMARIA DA GRAC¸A TEMIDODedicated to Professor Maria de Nazar´e Mendes LopesAbstract. In this work a stationary standard gaussian random field {Xn,m,(n,m) 2 N2}, with correlations rn,m satisfying rn,m ln(nm) !     0, n, m ! +1, is considered. We prove that the limit in distribution of the sequence of point processes of exceedances, for some normalized level, is a Cox process. Then, the maximum of n ⇥ m variables of the random field converges in distribution to a convolution of the Gumbel with the gaussian distribution. Our results extend the ones presented by Choi [2], where   = 0 is considered, as well as the results of Leadbetter, Lindgren, and Rootz´en [5], where the setup of real gaussian sequences is regarded.Pour m’avoir aid´e Une amie De m’avoir enseign´e Merci1. IntroductionExtreme value theory for random fields has been recently object of inten- sive research. Without being exhaustive, we mention the works of Choi [2], of Pereira and Ferreira [7], of Pereira [6], and references therein. In fact, a several results of the extreme values theory in the real line, that is for sequences of real random variables, stationary or not, have been extended when a random field is regarded. It should be pointed out that, traditionally, extreme values the- ory of gaussian processes has always received many esteem, being many timesAccepted: 14 February 2015.2010 Mathematics Subject Classification. 60G70, 60G60, 60G55, 60F05.Key words and phrases. Extremes, limit in distribution, point process, gaussian randomfields.This work was partially supported by the Centre for Mathematics of the Univer-sity of Coimbra – UID/MAT/00324/2013, funded by the Portuguese Government through FCT/MEC and co-funded by the European Regional Development Fund through the Part- nership Agreement PT2020.123]]></page><page Index="134" isMAC="true"><![CDATA[124 M. G. TEMIDOthe pioneers when some new theory appears, due both to its mathematical tractability either by its numerous applications.Let {Xn,m} be a stationary standard gaussian random field on N2, with correlations ri,j = E(X`,mX`+i,m+j), such thatri,j = r|i|,|j| (1.1)for each i and j in Z. We note that this condition is satisfied if the random field is isotropic, that is, if ri,j, for (i,j) 2 Z2, is a function only of the euclidean norm k(i, j)k (cf. Adler [1]).Let {Nn,m} be the sequence, doubly indexed, of the point process of ex- ceedances of a real high level un,m by the random variables X1,1, . . . , Xn,m, defined byNn,m(B)=]⇢(i,j)2N20 :✓i,j◆2BandXi,j>un,m , 8B✓]0,1]2. (1.2) nmIn this paper we study the convergence in distribution of the sequence of point processes of exceedances, {Nn,m}, when the correlations of the random field {Xn,m} satisfy a strong dependence condition, specified byrn,m ln(nm) !     0, n,m ! +1. (1.3)The case   = 0 is studied in [2] for stationary gaussian random fields and in [6] for non stationary ones. In fact, considering that {Xn,m} is a non stationary gaussian random field, in [6] is studied its extremal behaviour, assuming that the correlations r(i1 ,i2 ),(j1 ,j2 ) satisfy |r(i1 ,i2 ),(j1 ,j2 ) | < ⇢|i1  j1 |,|i2  j2 | , for some sequence {⇢n,m} such that ⇢n,m ln(nm) ! 0, n,m ! +1, ⇢n,0 lnn ! 0, n ! +1and⇢0,mlnm!0, m!+1.We end this section recalling the underlying main results for gaussian se- quences of real random variables.Throughout this work   and   denote, respectively, the standard gaussian distribution function and its density.Consider that {Xn} is a stationary standard gaussian sequence, with cor- relations sequence {rn } satisfying rn log n !     0, n ! +1. The sequence of point process of exceedances of the high level un by X1, X2, . . . , Xn is given by Nn(B)=#{j 1:j/n2B,Xj >un},foreachBorelsubsetBof]0,1].Itis well known that (cf. Leadbetter, Lindgren, and Rootz´en [5]), if un is such that n(1  (un)) ! e x, n ! +1, x 2 R, the sequence {Nn} converges in distribu- tion to a Cox process with stochastic intensity exp( x  +p2 Y ), where Y is a standard gaussian random variable. If   = 0, the limit in distribution of {Nn} is, obviously, a Poisson process with intensity   = e x. As a trivial corollary, the sequence of maxima {Mn }, with Mn = max{X1 , . . . , Xn }, under the clas- sicalnormalizationun :=un(x)=x/p2lnn+p2lnn ln(4⇡lnn)/(2p2lnn),]]></page><page Index="135" isMAC="true"><![CDATA[STONGLY DEPENDENT GAUSSIAN RANDOM FIELDS 125converges in distribution to a random variable whose distribution is a mixture of Gumbel distributions regulated by Y . For   = 0, the classical Gumbel limit is obtained. In the scope of non stationary gaussian sequence, extensions of these results can be found in Hu¨sler [3] and in Temido [8].2. Main lemmasWhen dealing with the distribution function of gaussian vectors, the Normal Comparison Lemma ([5]) appears as a widely useful result. This lemma esta- blishes bounds for the di↵erence between two multivariate gaussian distribution functions, which are convenient functions of their covariances. Our next lemma is a particular version of that result, with a specific approach for the random fields.Lemma 2.1. Let {Xn,m} and {Yn,m} be stationary standard gaussian randomfields with correlations ri,j and vi,j, respectively, satisfying (1.1). Consider thatun,m,n,m 1arerealnumbers.Then,forQn,m ={(i,j):1in,1j(2.1)m}andIn,m ={(i,j):(i,j)6=(0,0),0in,0jm},itholdsthat  0 10 1   \ X \ !   P @i,j 2Qn,m{Xi,j  un,m}A   P @i,j 2Qn,m{Yi,j  un,m}A Knmwith K > 0 and wi,j = max{|ri,j |, |vi,j |}.(i,j )2In,m|ri,j  vi,j|exp   u2n,m , 1+wi,jIn order to specify an approximation for the distribution of the vector {Xi,j : (i, j ) 2 Qn,m } of the gaussian random field, in the next lemma we give conditions which imply that the upper bounds in (2.1) are asymptotically zero. This lemma is an extension of Lemma 5.2.1 of [2], where only the assump- tion rn,m ln(nm) ! 0, n, m ! +1, is considered. It should be remarked that suitable conditions on the behaviour of the sequences {r0,j}j and {ri,0}i are also needed.Lemma 2.2. Consider that {rn,m} and {un,m} are real sequences,   is a non negative real number and ⇢n,m =  /ln(n m). Assume that {n m(1  (un,m ))}n,m , {rn,0 ln n}n Xand {r0,m ln m}m are bounded sequ!ences and that (1.3) holds. Thennm |ri,j  ⇢n,m|exp   u2n,m !0, n,m!+1, (2.2)1+wi,jwith wi,j = max{|ri,j|,⇢n,m} and In,m defined in Lemma 2.1.(i,j )2In,m]]></page><page Index="136" isMAC="true"><![CDATA[126 M. G. TEMIDOProof. We first assume   > 0. When   = 0, the proof of (2.2) for (i,j) in In,m\{(0, 1), (1, 0)} can be seen in [2]. By now we also assume that {un,m} is a sequence of normalized levels, that isnm(1  (un,m))!⌧, n,m!+1.Consider  (k, `) = sup wi,j and let ↵ be a real number such thati k,j `0<↵< 1  (1,0)  1  (1,1).1+ (1,0) 1+ (1,1)Take also p := [n↵], q = [m↵],   :=  (1,1) and  ⇤ :=  (1,0).We split the sum in (2.2) into eight terms. Namely, with A1 := {i 2 N : 1ip},A2 :={i2N:p<in},B1 :={j2N:1jq}, B2 :={j2N:q<j<m},weconsiderthesubsetsA1⇥B1,A1⇥B2,A2 ⇥B1, A2 ⇥B2, A1 ⇥{0}, A2 ⇥{0}, {0}⇥{B1} and {0}⇥{B2}.In this proof K, K1, K2,... denote suitable constants.Take into account that 1  (x) ⇠  (x)/x, x!+1, from nm(1  (un,m))!⌧, n, m ! +1,!we deduce exp  u2n,m ⇠Kun,m,2 n m that is un,m ⇠ p2 ln(n m),K >0,and 2ln(nm) ⇠1, n,m!+1, u 2n , m! +1. Considering un,m = x/an,m +Xp Xq u2n,m nm |ri,j  ⇢n,m|exp  1+wi,ji=1 j=1 2n1+↵m1+↵ exp⇣ u2n,m ⌘1+ ⇣ ⇣u2 ⌘⌘2  2n1+↵m1+↵ exp   n,m 1+ 2Kn1+↵ 2 m1+↵ 2 (u )2 1+  1+  n,m 1+ 1 1+  1+ K (nm)1+↵  2 (ln(nm)) 1  !0, n,m!+1,n, m bn,m, x 2 R, similarly, we obtaina = (2 ln nm)1/2 and b n,m n,m  a 1 ln(4⇡ ln(n m))/2. (2.3) n,m n,m= aFor the term concerning A1 ⇥ B1, we have s!uccessively]]></page><page Index="137" isMAC="true"><![CDATA[STONGLY DEPENDENT GAUSSIAN RANDOM FIELDS because1+↵< 2 .WhenA2⇥B1 isconsidere!d,weget1+ 127nq2 XX ! un,mnm |ri,j  ⇢n,m|exp  1+wi,j i=p+1 j=1u2n,m XnXqnmexp  1+ (p,1) i=p+1j=1|ri,j  ⇢n,m|u2 nm exp  n,m!! 2 1+ (p,1)Xn Xq ⇢     +         +           (2.4)Now, since lnn+lnm  lnn⇥lnm, for n,m   7, and  (p,1)  k , with lnpk > 0, we obtain, for n and m large enough, ri,j   ln(ij) ln(nm↵)   ln(nm↵) ln(nm) 2⇥i=p+1 j=1ln(ij)" !#2n2m u2 1+ (p,1) n2m ln p 2 ln pexp   n,m  Klnn 2 lnn(ln(nm))k+ln p (nm) k+ln pK 1 n2  2↵lnn m1  2↵lnn (lnn+lnm) ↵lnn1 k+↵ ln n k+↵ ln n k+↵ ln n lnn1 k+↵lnn k+↵lnn Kn 2k (lnn)  k(lnm)↵lnn m1 2 ↵lnn k+↵lnn k+↵lnn(2.5)exp⇣ 2klnn ⌘⇣k+↵lnn⌘(lnm) ↵lnn m1 2 ↵lnn=KK (lnm) ↵lnn m1 2 ↵lnn .12 k+↵ ln n k+↵ ln nk+↵ ln n k+↵ ln nexp klnlnn k+↵lnnFurthermore, it can be shown thatlnn Xn Xq  ri,j        lnn Xn Xq 1 |ri,j ln(ij)  |n i=p+1 j=1 ln(i j) n i=p+1 j=1 ln(i j) 1 m↵ sup |ri,j ln(i j)    |. (2.6)k+↵lnn1+↵  2↵ ln n < 0, for n large, for the first sum in the right hand side of (2.4), k+↵lnn↵ i 1,j  1Thus, due to (2.5) and (2.6) and taking into account that ↵ ln n < 1 and]]></page><page Index="138" isMAC="true"><![CDATA[128we deducen 2 m lnnOn the other hand, we haveu 2 exp  n,mM. G. TEMIDO!!2    1 +   ( p , 1 ) l n n Xn Xq       ri,j    n i=p+1j=1 ln(ij)2k+↵ ln n k+↵ ln n i, jK(lnm) ↵lnn m1+↵  2↵lnn sup |r ln(ij)  |i 1,j  1 k+↵ ln n i, jKlnmm1+↵  2↵lnn sup |r ln(ij)  |!0, n,m!+1. i 1,j  1u2 nm exp(  n,m)!2   1+ (p,1) Xn Xq            ↵   i=p+1j=1 ln(ij) ln(nm ) K  K1 ln ↵  ↵ ↵lnni=p+1j=1 nm nmn 2 m 1 + ↵ ↵u 2 exp(  n,m)!2  1 +   ( p , 1 )   Xn Xq   i j   12ln(nm ) n2m1+↵2u2 exp(  n,m)! 2↵ ln n k+↵ ln n  Z 1 Z 1 ↵lnn0 0|lnxy|dxdylnn (lnn)  k2k+↵ ln n  2klnm2↵ ln n1nk+↵ ln n exp⇣ klnlnn⌘m 1 ↵+k+↵ ln n lnn⇣k+↵lnn⌘ lnm 1 !0, n,m!+1. 2↵ ln n=KSimilarly, for the third sum of the right hand side of (2.4), we obtain1exp  2klnn m 1 ↵+k+↵lnn lnn k+↵lnnu2 nm exp   n,m!!2     1+ (p,1) Xn Xq          ↵     i=p+1j=1 ln(nm ) ln(nm)2◆lnp ✓ ln(nm) k+lnp    lnn  +   lnn  ✓lnn (nm)2  ln(nm↵)   ln(nm) For the sum concerning A1 ⇥ B2 the desired result is obtained changing n with m.K n2m1+↵1 k+↵ ln n    ◆K ,lnmm1+↵  2↵lnn !0, n,m!+1.]]></page><page Index="139" isMAC="true"><![CDATA[STONGLY DEPENDENT GAUSSIAN RANDOM FIELDS 129 Likewise, we deal with the sums related to A2 ⇥ B2. Indeed, since (p,q)ln(nm)  k, with k > 0, we getn m Xn Xm         r i , j             e x p   u 2n , m ! i=p+1 j=q+1 ln(n m) 1 + wi,j!! 2n2m2 u2 1+ (p,q) ln(n m)exp   n,mln(nm) 2 nmK n2m2 ln(nm) k+↵2ln(nm):i>p,j>q i,j ln(n m) i=p+1 j=q+1   nm   n m ;⇥ 8< Xn Xm       r           +         :i=p+1j=q+1  i,j ln(ij)   ln(ij)      9= ln(nm) ;   ⇥ 8< s u p | r l n ( i j )     | +   Xn Xm       l n i j       1 9=✓ ◆ ↵2 ln(n m) 1 ln(n m) (n m)2deducei 1 u2n,m !nmXp i=1|ri,0  ⇢n,m|exp 1+wi,0exp⇣ klnln(nm)⌘ ⇢ ⇣k+↵2 ln(nm)⌘⇥ o(1)+ Z 1Z 1 exp  2kln(nm) ln(nm) 0 0 |lnxy|dxdyK2K3{o(1)+o(1)}=o(1), n,m!+1.k+↵2 ln(n m)Let us consider now the sums associated with A1 ⇥ {0} and A2 ⇥ {0}. From nowonwetake  0.Duetothefactthatsupwi,0  sup wi,j = ⇤,wei 1,j  0 2n1+↵mKn1+↵  2⇤m1  2⇤ (ln(nm)) 1⇤1+ K(nm)1+↵  2 ⇤ ln(nm)!0, n,m!+1.!! 2 ⇤ u2 1+ exp  n,m 21+  1+  1+ ]]></page><page Index="140" isMAC="true"><![CDATA[130 M. G. TEMIDOWrite ⇤⇤ :=supwi,0.Sincethesequence{ ⇤⇤lnn}isboundedand ⇤⇤lnnki pimplies  ⇤⇤ < 1 for n large, we obtain2nmu2n,m ! |ri,0  ⇢n,m|exp   !Xni=p+1 1 + wi,02 ⇤⇤ u2n,m 2nm  exp  1+ ⇤⇤(n m)2= K  ⇤⇤ ln n +  ⇤⇤ ln m exp (2 ⇤⇤ ln n)! 2 n2 m  ⇤⇤ exp   u2n,m  exp K n2m  ⇤⇤ ln(n m) exp (2  ⇤⇤ ln n + 2 ⇤⇤ ln m)u2n,m ⇤⇤ 1+ ⇤⇤✓ m1 2 ⇤⇤ ◆K 1+lnm!0,m!+1.1 m1 2 ⇤⇤ m1 2 ⇤⇤For the sums concerned the set {0} ⇥ (B1 [ B2) we use the same arguments. Suppose now that {un,m} is a real sequence satisfying n m(1    (un,m))  ⌧ and that {vn,m} is another real sequence such that n m(1    (vn,m)) ! ⌧, asn,m!+1.Thus,formandnwithpn2 +m2>M,withMlargeenough, we get vn,m  un,m. Then, since (2.2) holds for {vn,m} it follows the same for{un,m }.⇤3. Point processes of exceedancesConsider the Cox process N defined over ]0,1]2, with stochastic intensity givenbyexp  x  +p2 Y ,wherex2R, 2R+0 andY isastandard gaussian random variable. The distribution of N is defined, in terms of its probability mass function, byP\q ! N(Bi) = sii=1Z +1 Yq  (Bi)e x  +zp2 = s !  1 i=1 i! x  +zp2 ⇥ exp(  (Bi)e )  (z)dz(3.1)where, for q   1, s1,...,sq are non negative integers, B1,··· ,Bq are disjoint subsets of ]0, 1]2 and   represents the Lebesgue measure.]]></page><page Index="141" isMAC="true"><![CDATA[STONGLY DEPENDENT GAUSSIAN RANDOM FIELDS 131The following theorem, the main result of this work, states that a Cox process is obtained as the limit in distribution of the sequence of point processes of the exceedances defined by (1.2). Is then established the expected extension of the classical results for gaussian sequences presented in [5], when     0 is regarded, as well as the ones for random fields due to [2] for   = 0.Theorem 3.1. Let {Xn,m} be a stationary standard gaussian random field with correlations satisfying (1.1) and (1.3) and suppose that the sequences {rn,0 ln n} and {r0,m ln m} are bounded. Consider un,m = x/an,m + bn,m with x 2 R and an,m and bn,m defined by (2.3). Then the sequence {Nn,m} of point processes of exceedances of the high level un,m by the random field {Xn,m} converges in dis- tribution to the Cox process with stochastic intensity exp   x     + p2  Y  , that is, the point process with distribution characterized by (3.1).Proof. We use a Kallenberg’s theorem (see [4, Theorem 4.7] or [5, Theo- rem A.1]). Since for each ]a, b]⇥]c, d] ✓ ]0, 1]2 we getE (Nn,m (]a, b]⇥]c, d])) = n(b   a)m(d   c) (1    (un,m )) ! (b a)(d c)e x, n,m!+1,it remains to prove thatP ({Nn,m(B) = 0}) ! P ({N(B) = 0}), n,m ! +1,for all B of the form Sq`=1 B`, with q 2 N and B1,...,Bq disjoint subsets of ]0, 1]2.We use the notationsMn,m (]a,b]⇥]c,d]) = max{Xi,j : [an] < i  [bn] , [cn] < j  [dn]}andMn,m (]a,b]⇥]c,d],⇢n,m) = max{Yi,j : [an] < i  [bn] , [cn] < j  [dn]}where {Yi,j} represents a standard gaussian random field with all the covari- ances equal to ⇢n,m =  / ln(n m) (except ⇢0,0 = 1).Let B1, · · · , Bq be disjoint subsets of ]0, 1]2 and let Y be a standard gaussian randpom variable independent of {Mn,m(Bi,0),i=1,2,···q}. Consequently, for each n and m and for each B`, the random variables Mn,m (B`, ⇢n,m) and 1   ⇢n,mMn,m (B`, 0) + p⇢n,mY have the same distribution. Then, with]]></page><page Index="142" isMAC="true"><![CDATA[132 M. G. TEMIDOwn,m = p1   ⇢n,m(un,m   zp⇢n,m) we obtainP =P\q ! {Mn,m (B`, ⇢n,m)  un,m}`=1 !\q  p\1 ⇢n,mMn,m(B`,0)+p⇢n,m Y un,m`=1 Z+1 q=Moreover, due to Lemma 2.2, we get  \q ! \q  P {Mn,m(B`)  un,m}   P`=1 `=1! 1`=1P{Mn,m(B`, 0)  wn,m}  (z)dz.(3.2)when n,m ! 1.Let {Xn⇤,m} be a standard gaussian random field of independent randomvariables and denote by {Nn⇤,m} the sequence of point process of exceedances, of a normalized level, by the random variables X1⇤,1, . . . , Xn⇤,m.Mutatis mutandis, we can establish the expected generalization of Corol- lary 5.2.2 of [5]. Indeed, for each B ⇢]0,1]2 and for all s 2 N0, we have⇤ e ⌧ (B)(⌧  (B))sP(Nn,m(B) = s) ! s! , n,m ! +1, (3.3)where ⌧ = lim nm(1    (x/an,m + bn,m)). Furthermore, for disjoint sub- n,m!+1    sets of ]0, 1]2, B1, · · · , Bk, the vector Nn⇤,m(B1), · · · , Nn⇤,m(Bk) , converges in distribution to the product of limits in (3.3).!  {Mn,m(B`, ⇢n,m)  un,m}   ! 0,Now, taking into account that p1 ⇢n,m(un,m zp⇢n,m)= 1 1⇢n,m+on,m(⇢n,m) ✓ x +bn,m zp⇢n,m◆2 an,m = x+  zp2  +b +o(a 1 )an,m n,m n,mand then ⌧ := ⌧(x) = x + ⌧   zp2 , it follows that]]></page><page Index="143" isMAC="true"><![CDATA[`=1STONGLY DEPENDENT GAUSSIAN RANDOM FIELDS133q! \\pp!P{Mn,m(B`, 0)  1   ⇢n,m(un,m   z ⇢n,m)} q x +     zp2 ⇠ P {Mn,m(B`, 0)  a + bn,m} `=1 ! n,m\q Yq p = P {Nn⇤,m(B`) = 0} ! exp( ( (B`)e x  +z2  ),n, m ! +1.`=1 `=1Finally, by dominated convergence and due to (3.2), we deduce\q !\q !! P {Nn,m(B`) = 0} = P {Mn,m(B`)  un,m}`=1Z + 1 Yq ⇣   x     + z p 2   ⌘ \q! exp   (B`)e  (z)dz = P {N(B`) = 0}  1 `=1 `=1`=1as n, m ! +1.As a corollary we obtain the convergenceP (Mn,m  un,m) ! Z +1 exp ⇣ e x  +zp2  ⌘  (z)dz,  1that generalizes the well known univariate result ([5]).References⇤[1] R. Adler, The geometry of random field, John Wiley and Sons, New York, 1981.[2] H. Choi, Central limit theory and extremes of random fields. PhD Thesis, Univ. of NorthCarolina at Chapel Hill, 2002.[3] J. Hu¨sler, Asymptotic approximation of crossing probabilities of random sequences, Z.Wahrsch. Verw. Gebiete 63 (2), 257–270, 1983.[4] O. Kallenberg, Random measures, Academic Press, London-New York, 1976.[5] M. R. Leadbetter, G. Lindgren, and H. Rootz´en, Extremes and Related Properties ofRandom Sequences and Processes, Springer-Verlag, Berlin, 1983.[6] L. Pereira, On the extremal behavior of a nonstationary normal random field, J. Statist.Plann. Inference 140 (11), 3567–3576, 2010.[7] L. Pereira and H. Ferreira, Point processes of exceedances by random fields, J. Statist.Plann. Inference 142 (3), 773–779, 2012.[8] M. G. Temido, Mixture results for extremal behaviour of strongly dependent nonstatio-nary Gaussian sequences, Test (9), 439–453, 2000.(M. G. Temido) Department of Mathematics, University of Coimbra, Apartado 3008 EC Santa Cruz, 3001-501 Coimbra, PortugalE-mail address: mgtm@mat.uc.ptn, m ! +1,]]></page><page Index="144" isMAC="true"><![CDATA[]]></page><page Index="145" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING APPROXIMATE BAYESIAN COMPUTATIONP. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMANDedicated to Nazar´e LopesAbstract. Bilinear processes are highly flexible for modeling nonlinear and heavy-tailed features that are often exhibited by financial and envi- ronmental time series. However, intractable likelihood functions, together with the lack of verifiable conditions of invertibility and stationarity, ex- cept for very simple bilinear processes, constraint their use as models. The traditional methods of least squares and conditional maximum like- lihood (CML) do not give satisfactory results, particularly for heavy- tailed series. This paper aims to show the advantages and drawbacks of using simulation-based estimation techniques for bilinear models. Ap- proximate Bayesian Computation (ABC), as well as sequential Markov chain Monte Carlo (MCMC) methods are often applied with success as inferential tools for time series. Due to ease with which one can simulate bilinear processes, ABC is particularly a promising inferential alterna- tive. We assess the viability of ABC as an inferential tool for the bilinear processes. The performance of the ABC algorithm for parameter esti- mation is assessed using several simulated samples from simple bilinear models. The results are compared with the parameter estimates obtained using the CML method.1. IntroductionToday, we are more aware that many observed time series, particularly those coming from financial markets and environmental sciences, do not conform with the traditional assumptions of linearity and Gaussian error structures, often exhibiting heavy-tailed behaviour. Therefore, models usually applied are increasingly more complex, capturing such features as nonlinear variations bothAccepted: 11 February 2015.2010 Mathematics Subject Classification. 62M10, 91B84.Key words and phrases. Bilinear processes, ABC, conditional maximum likelihood, se-quential MCMC.The work was financially supported by Funda¸ca˜o para a Ciˆencia e a Tecnologia (FCT),Portugal, through the projects PEst-OE/MAT/UI0006/2014 and PTDC/MAT/118335/2010.135]]></page><page Index="146" isMAC="true"><![CDATA[136 P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMANin the mean and variance. Although some of these nonlinear classes of models, such as the GARCH are analytically tractable, likelihood functions for the class of bilinear processes are analytically intractable. For heavy-tailed bilinear processes, least square methods do not give satisfactory results and therefore, inference for such processes can be a challenging task.The use of MCMC methods in a Bayesian context has proven to be a very powerful solution. The seminal paper by Gelfand and Smith [21] established the use of simulation-based approaches in a Bayesian framework. The use of the Gibbs Sampler, introduced by Geman and Geman [22], become a common way to overcome the di culties of handling complicated posterior distributions. After those ground-breaking papers, many other algorithms were developed to deal with specific problems. The Metropolis Within Gibbs, proposed by Gilks et al. [23] for addressing the simulation from non-logconcave full conditional distributions, the Slice Sampler (Neal [28]) or the Reversible Jump, proposed by Green [16] in order to handle estimation problems associated to models with varying-dimension parameter spaces, are such examples. The above mentioned MCMC methods depend heavily on the concept of likelihood and the existence of analytical expression for the likelihood. The Sequential Monte Carlo (SMC) in general, and Particle Markov chain Monte Carlo in particular, are extensions that enable dealing with complex likelihoods sequentially, bringing some nu- merical ease into the inference. In contrast, the ABC algorithms are suggested for handling situations when it is not possible to give any tractable analytical expression for the likelihood.The purpose of this paper is to assess the viability of using an ABC algorithm for estimating the parameters of simple bilinear processes.The structure of the paper is as follows. The bilinear models are reviewed in Section 2. The ABC algorithms are introduced in Section 3. Some simulation results are provided in Section 4 and finally comments and conclusions are given in Section 5.2. The Bilinear models2.1. Preliminaries. Bilinear models (Subba Rao and Gabr [36]) are possibly the most natural way to extend the ARMA models for explaining features such as nonlinearity and heavy-tailed behavior. Yt is called a bilinear process, BL(p, q, m, k), if it satisfies the di↵erence equation:Xp Xq XmXkYt   ajYt j = cj✏t j + b`1`2 Yt `1 ✏t `2 . (2.1)j=1 j=0 `1=1 `2=1]]></page><page Index="147" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 137Here, c0 = 1 and {✏t} is a sequence of independent and identically distributed (iid) random variables (rvs) with zero mean and variance  2.In this class, the conditional mean E(Yt|Ft 1) is a nonlinear function of yt i,i = 1,2,.., whereas the conditional variance Var(Yt|Ft 1) is constant. Here Ft 1 is the  -field generated by (Yt 1, Yt 2, . . . ). Further extension of this class can be made by including terms indexed to `2 = 0 in the last summation of (2.1), in which case these models also account for nonlinear variations in the conditional variance.The class of bilinear models plays an important role in modeling nonlinearity for various reasons.(1) The class is an obvious generalization of ARMA models resulting in nonlinear conditional mean and conditional variance.(2) Under fairly general conditions, bilinear processes approximate finite order Volterra series expansions to any desired order of accuracy over finite time intervals (see Brockett [4]). Volterra series expansion is a dense class within the class of nonlinear time series. Therefore, under fairly general conditions, bilinear processes are also a dense class within nonlinear processes, approximating any nonlinear process to a desired level of accuracy.(3) The class is fairly well-studied. Much is known regarding the existence of unique and stationary solutions. Although some identification, esti- mation and diagnostic techniques are available, much of the work on the class remains to be completed.(4) Bilinear processes are often used in control theory although in a some- what di↵erent context in which they are applied in time series. In con- trol theory, the output Yt, as well as the input processes ✏t are observ- able, making the probabilistic structure simple. In the context we use these models, the input random process ✏t is not observed. This some- what restricts the use of these models within the time series framework. In estimation and prediction, it is important to know that the input process ✏t is also measurable with respect to the Ys, s  t, i.e., it is invertible. Unfortunately, the lack of verifiable conditions for invertibil- ity (except for very simple bilinear processes) restricts the use of these processes as models.(5) A very important feature of the bilinear processes is that they are capable of producing sudden bursts of large values. Hence, they are very appropriate for modeling time series showing burst-like phenomena. This behavior is common in many areas, such as telecommunications and internet tra c (see Resnick [31] and Turkman et al. [38]).]]></page><page Index="148" isMAC="true"><![CDATA[138P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMAN(6) It is known that the heavy-tailed behavior of bilinear models is a con- sequence of the tail weight of the innovation process, as well as the multiplicative, nonlinear, features of the di↵erence equation given in (2.1). For instance, when the innovations ✏t are regularly varying, the tail of the bilinear model BL(0, 0, 1, 1),Xt = cXt 1✏t 1 + ✏t,where c is a constant, satisfies the following relationP(|Xt |>x)⇡P(✏21 >x). For a general BL(p, q, m, k), we haveP(|X |>x)⇡P(✏k+1 >x). t1Hence, while for linear models, Xt and ✏t are tail equivalent, for BL(p, q, m, k) models X is tail equivalent to ✏k+1 (see Resnick [31]).ttWhen innovations have lighter tails, for example Gaussian, then the relationship between the tails of input-output series is more compli- cated. However, typically the tail of the output series Xt is regularly varying when the innovations are Gaussian (Turkman and Amaral Turkman [40]). These results clearly indicate why and how bilinear processes produce extreme observations and show how useful they can be in modelling heavy-tailed dependent data.These features are evident in the simulated data sets that are plotted in Figure 1. When the innovation process is Gaussian (light-tailed), the simulated sample does not exhibit very large values (left panel). On the other hand, whenXt−5 0 5Xt0 50 100 150 200 250 300 350Xt0 10000 20000 30000 40000 500000 100 200 300 400 500 0 500 1000 1500 2000 2500 3000 0 500 1000 1500 2000 2500 3000 tttFigure 1. Simulated bilinear data sets - N(0,1) innovations (left), Pareto(2.5) innovations (center) and Pareto(1.5) innovations (right).]]></page><page Index="149" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 139the errors are Pareto(↵), ↵ > 0, with ↵ = 2.5 and ↵ = 1.5, the process can generate very large values (center and right panels). It is known that, for ↵ = 2.5, the Pareto distribution, with distribution function given by F (x) = 1 x ↵, x > 1 and ↵ > 0, has both finite mean and variance, but when ↵ = 1.5, although the mean is finite, the variance is not. The situation can be more extreme if ↵  1 because neither the mean nor the variance are finite.As mentioned before, the conditions that guarantee invertibility and sta- tionarity of bilinear process can only be derived for some particular models (see Subba Rao and Gabr [36]), such as the simple bilinear process which has only one bilinear term. Invertibility is a fundamental requirement if the model is to be used for prediction. Let Yt be a simple first order bilinear process, BL(1,0,1,1), given byYt = aYt 1 + bYt 1✏t 1 + ✏t, (2.2)where {✏t} are iid rvs. A su cient condition for the process to be stationarity is given by Pham and Tran [30]: the parameters a, b and  2 must satisfy the condition a2 + b2 2 < 1, where |a| < 1. A su cient condition for invertibility is that the parameters a and b are such that b2E(Yt2)  1, as proved by Pham and Tran [30]. They also show that, provided the {et} are Gaussian rvs, the condition2(1+a)b4 4 +2(1 a)b2 2  (1 a)2(1+a)0 (2.3)is su cient for the simple bilinear process to be invertible. Additional condi- tions on a, b and  2 are required if we want the bilinear process to have up to the fourth moment finite. In this case, besides | a |< 1 and a2 + b2 2 < 1, the following (necessary) conditions | a3 +3ab2 2 |< 1 and a4 +6a2b2 2 +3b4 4 < 1 have to be satisfied (Sesay and Subba Rao [34], Kim et al. [24], as referred by Leon-Gonzalez and Yang [25]).All these conditions severely restrict the parameter space while searching for admissible estimators for the parameters.Although bilinear processes do not satisfy the Markov property, it is possible to write them as state space models (SSM) satisfying a Markovian property, which may be very helpful in several aspects, namely for using SMC algo- rithms. If the observable process Yt is a BL(1,0,1,1) given in (2.2) then, solving iteratively for Yt, we get, after n iterations,n n 1 " j #Yt =Y(a+b✏t i)Yt n +X Y(a+b✏t i) ✏t j +✏t.i=1 j=1 i=1 If Xt = (a + b✏t)Yt, thenXt = (a + b✏t)Xt 1 + (a + b✏t)✏t]]></page><page Index="150" isMAC="true"><![CDATA[140 P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMANandYt = Xt 1 + ✏t.Here, Xt is a Markov process and Yt has the standard latent Markov processrepresentation.2.2. A brief review of parameter estimation. Traditionally, the bilinear parameter estimation is carried by the CML method for BL(p, 0, p, k) models (Subba Rao [35]). The process involves using some iterative method, such as the Newton-Raphson that requires the choice of initial values for the parameters close to the true values, which may not be an easy task in applications. The convergence of the algorithm may be very di cult to achieve due (even for the simple bilinear model) to the invertibility condition that constraints the values of a and b. To reduce this problem, Scotto [33] imposed some restrictions to the maximization algorithm by using some penalty functions. The method of least squares for the model (2.1) with Gaussian innovations works reasonably well (Subba Rao [35]) and is equivalent to the method of estimating functions (Turkman et al. [38]).Grahn [15] developed a conditional least squares (CLS) approach for estimating the parameters of a superdiagonal bilinear model and also for a standardized version of the BL(p,0,p,1) model. Several other methods, such as the use of Yule-Walker type di↵erence equations for higher order cumulants (BL(p, 0, p, 1) model), the Extended Least Squares, the Recursive Prediction Error Method and the Extended Kalman filter (BL(p, 0, p, 1) model) have also been used for estimating the parameters of particular bilinear models (see Subba Rao and Silva [37], Gabr [19, 20]). Frequency-domain methods for estimating the param- eters of the BL(p, 0, p, 1) model have also been used by Sesay and Subba Rao [34]. Feng et al. [14] considered the model Yt = b✏t 1Yt 2 + ✏t, t = 1,2,...,n, with {✏t} iid N(0,  2) and used empirical likelihood (EL) in order to obtain con- fidence intervals for the parameter b. Quite recently, Ling et al. [26] proposed a GARCH-type maximum likelihood estimator (MLE) for the parameters of the bilinear model Yt = μ + aYt 2 + bYt 2✏t 1 + ✏t, where {✏t} are rvs iid with zero mean.3. The ABC and the SMC algorithmsSMC methods enable processing the data in blocks, as the observation be- came available, thus diminishing the computational burden of complex likeli- hoods. They are particularly useful for dealing with dependent data, such as time series. The volatility of stock exchange, the movements of airplanes cap- tured by radars and industrial production are situations in which the processing]]></page><page Index="151" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 141of data is naturally done sequentially. In these cases, decisions have to be taken immediately as data becomes available and conditionally on the existing infor- mation. The decisions cannot be postponed until the entire data set becomes available (Doucet et al. [13]). In this framework of bayesian inference, and con- sidering that the focus is on the posterior and/or the predictive distributions, SMC methods simulate values sequentially from the intermediate distributions, in the sense that only the available data (up to the moment) are used.ABC algorithms appear in the literature as a way to overcome computational problems associated to models that have analytically complicated or intractable likelihoods. In this framework, using MCMC methods can be very hard or even impossible to accomplish. The ABC algorithms are based on simulated values from the model of interest for a given set of parameter values and the information contained in the sample simulated is resumed in a set of summary statistics which is compared to the same summary statistics obtained from the fixed observed sample. The set of parameter values is then kept or discarded according to how these sets of summary statistics relate. The success of the ABC algorithm that is implemented depends on several factors, such as how easy it is to simulate data from the model of interest and how accurately do the summary statistics represent the data in hand. In this bilinear framework, there is an obvious lack of summary statistics. Although the Bayesian approaches via ABC method were developed in a population genetics context, they are nowadays used in several other areas such as archeology, epidemiology and ecology (see, for instance, Beaumont et al. [1], Csill´ery et al. [9]) and Marin et al. [27] for further references).3.1. The ABC algorithm. On what follows, let ✓ be a parameter, or a vector of parameters, associated to a sampling model M, p(✓) the prior distribution of ✓ and Dt = (x1,x2,...,xt) the observations available at instant t (t   1). The posterior distribution of ✓ is referred as p(✓ | Dt) / L(✓ | Dt)p(✓), where L(·) is the likelihood.Let us denote by Dt, the fixed observed data set. Similarly, denote by S, a vector of summary statistics which give a good representative of the model. Ideally, such vector of summary statistics should be su cient for the model, but when the likelihood is not known or is intractable, then su cient statistics other than the ordered sample will not exist. Let S0 = S(Dt) and S⇤ = S(Dt⇤) be respectively be the value of the summary statistics calculated at the fixed observed and simulated samples Dt and Dt⇤. The ABC algorithms are based on the widely know rejection method (Paulino et al. [29]). Essentially, they consist]]></page><page Index="152" isMAC="true"><![CDATA[142 P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMANon carrying out the following steps:Step 1. simulate a value of ✓ from the prior distribution, p(.)Step 2. simulate a new sample Dt⇤ from the model M corresponding to this parameter value.Step 3. accept the sampled value ✓ if d(S(Dt⇤), S0)   .Here, d(.,.) is a metric and   is a certain, nonnegative, tolerance level. The simulated data are observations from p(✓ | d(S(D⇤), S0)   ), but not from the true posteriori distribution, unless   = 0. If   ! 1 then the sampled ✓s are simply observations of the prior distribution, and not from the posteriori as it is intended (see Marjoram et al. [18] and Wilkinson [39]).Evidently, the performance of the algorithm depends on several factors, such as the distance function considered, the kind and the quality of the summary statistics used to summarize the data and the choice of  . The values of ✓ must balance the computational e ciency and the precision that we want the final results to have. The euclidean norm,vuXn d(x,y)=||x y||=t (xi  yi)2, x,y2Rn,i=1is the distance function usually considered.The number and kind of summary statistics depend on the problem in hand.The sample moments are natural choices. The first four moments compare the location, dispersion, asymmetry and kurtosis of the observed sample with the corresponding features of the data simulated from the model M, conditional to the value of ✓ simulated from the prior distribution. For instance, if the tail of the underlying distribution of an observed sample is heavy then a measure of tail weight similarity between the simulated and observed data should be contemplated. For example, the moments estimator of the tail index might be considered (Embrechts et al. [17]).In what concerns all these issues, there have been several developments in ABC methods in the last few years (see, for instance, Beaumont et al. [1] and Blum et al. [3]). Biau et al. [2] propose considering the kN values of ✓ that are closest to S0 for a certain proximity measure. Usually, kN corresponds to the 0.90 sample quantile. The ABC algorithms are easy to implement but are generally computationally intensive. For instance, according to Biau et al. [2], if the vector of parameters has dimension p = 3 and the number of summary statistics is m = 7, then N = 106 samples must be simulated in order to obtain a sample of size kN = 1000 of ✓, i.e.,p+4 kN ⇡Nm+p+4.]]></page><page Index="153" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 1433.2. The ABC algorithm for the simple bilinear process. For bilinear processes, no su cient statistics, other than the order statistics are known. Therefore, a set of summary statistics capturing several aspects of the model has to be chosen. The following seven summary statistics are selected in the case of the simple bilinear model that is being studied:(1) the median and the inter-quartile range;(2)  Y =h=11⇢2 2(h) where ⇢Y (h) and ⇢Y2(h) are hYYhYPP 1⇢2 (h) and  2 =1010h=1the values of the autocorrelation functions (ACF) of the time series{Yt} and of the squared values, {Yt2}, computed at lag h, respectively. The weight 1/h is used to give more emphasis to correlations associated to low lags.(3) The left and right tail indexes,  L and  R, using the moments estimator defined by Dekkers et al. [10] aswhereMn 1 kX  1 ˆ=M(1)+1 1 1 (Mn ) n 2 (2)M(j) =n k(logYn i:n logYn k:n)j," (1) 2# 1i=0where j = 1, 2. In this case, (Y1:n, Y2:n, . . . , Yn:n) represents an ascend- ing ordered sample of size n. This estimator can take any real value and as such is able to reflect heavy or light-tailed distributions, as well as exponential tails.(4) The right extremal index, ✓⇤ , where 0 < ✓⇤  1. The extremal index measures how strongly the largest values in the sample cluster. The smaller the value of the extremal index, the stronger the clustering of extreme observations is. If the largest values are independent then ✓⇤ = 1. However, the converse is not true (see, e.g., Coles [8] for details and examples). The usual estimate of the extremal index is ✓ˆ⇤ = nc/nu, where nu represents the number of exceedances over the threshold u and nc the number of clusters above u.The reason for choosing this set of statistics is to control the similarities between the first and second order properties, the degree of nonlinearity, as well as the tail behavior of the observed and the simulated time series, and the (possible) clustering of large values.]]></page><page Index="154" isMAC="true"><![CDATA[144 P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMAN4. Simulation study: ABC vs. CMLLet us consider the bilinear model given in (2.2). The innovations ✏t are iid rvs N (0, 1). A random sample of size n = 150 (Figure 2) is generated from the modelModel I: Xt = 0.2Xt 1 + 0.4Xt 1✏t 1 + ✏t.This model satisfies the condition a2 + b2 2 < 1 and also 2(1 + a)b4 4 + 2(1   a)b2 2   (1   a)2(1 + a)  0. The ABC results will be compared with the ones obtained by CML. The conditions a2 + b2 < 1 and 2(1 + a)b4 + 2(1   a)b2   (1   a)2(1 + a)  0 are incorporated in the maximization process just as Scotto [33] proposed.The initial estimate of a results from fitting an AR(1) model, as recom- mended in the literature. Instead of fitting a model to the residuals of the AR(1) model to obtain an initial estimate for b, as proposed by Subba Rao and Gabr [36], a grid of values for b was considered. The results are given in Table 1. The error considered in the Newton-Raphson algorithm was 0.0001. The models show that the estimation process seems to have converged to the solution (aˆ, ˆb) = (0.1796, 0.3224). The fastest result was attained with b0 = 0.3, although similar performances were obtained using the initial values b0 = 0.1, b0 = 0.2 and b0 = 0.4. It should be referred that, unexpectedly, the estimation of a was much more di cult than the estimation of b. Some of the models pre- sented in Table 1 occasionally “visited” regions of the parametric space where the condition |a| < 1 was not satisfied. However, the algorithm always found its way back to the “stationarity zone”. The estimate of  2 is quite good (⇡ 1.06) when compared to the true value.0 50 100 150ObservationsFigure 2. Fixed bilinear data sets - Model I (n = 150).y−3 −2 −1 0 1 2 3 4]]></page><page Index="155" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 145aˆ0 b0aˆ-0.6  0.1796147-0.5  0.1796168 -0.4  0.1796255 -0.3  0.1796203 -0.2  0.1796239 -0.1  0.1796172 0.1  0.179616 0.2  0.1796172 0.3  0.1796151 0.4  0.1796115 0.5  0.1796226 0.6  0.1796132 0.7  0.179616ˆb 0.3224413 0.3224415 0.3224426 0.3224419 0.3224424 0.3224416 0.3224413 0.322442 0.3224414 0.3224408 0.3224421 0.3224436 0.3224434 ˆ2 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738 1.059738Num. Iterations 281198775545812100.2214939anda⇠Unif[ 0.9,0.9], b⇠Unif([ 0.9, 0.1][[0.1,0.9])  2 |a,b⇠Unif[0.1,(1 a2)/b2],Table 1. CML estimates - Model I (n = 150) (The model did not converge for b0 =  0.7).As mentioned before, the bilinear models have likelihood functions di cult to handle, which makes the ABC methodology a especially attractive alternative for parameter estimation (Turkman et al. [38]).Let D be the fixed random sample of size n simulated from the model (2.2) with some a, b and  2. The prior distributions for (a,b, 2) are:where Unif[c,d] stands for the uniform distribution in the interval [c,d].The summary statistics to be considered are meant to reflect the similarities of location, dispersion, autocorrelation, tail-weight and the clustering of large values (not used in this case because the sample size is relatively small) between the fixed sample D and a very large number of samples (N = 107) simulated from the bilinear model presented in (2.2), conditional to the values a, b and  2. By considering the euclidean distance and the 0.90 percentile of the N = 107 distances calculated, a final sample of size kN = 1000 was obtained. The conditions that guarantee stationarity and invertibility were verified at eachiteration.The values of the approximate posterior distribution, representing the best1000 values of a, b and  2, corresponding to the sample of size n = 150 previ- ously simulated from the modelModel I: Xt = 0.2Xt 1 + 0.4Xt 1✏t 1 + ✏t,]]></page><page Index="156" isMAC="true"><![CDATA[146 P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMANare presented in Table 2. The results given by ABC for model I are not very good, specially for a. This situation is possibly due to the relatively low sample size.ABCMinimum Q0.25 Median Mean Q0.75 Maximum s------- 0.1089 0.4542 0.5729 0.5625 0.6749 0.8968 0.1523 0.3067 0.6143 0.7443 0.7802 0.9161 1.8460 0.2302Model I True values a=0.2 b=0.4  2 =1.0Table 2. Statistics for the sample of size kN = 1000 and the CML estimates of b and  2 obtained before.CMLaˆ=0.1796ˆb=0.3224  ˆ2 =1.0597Let us consider now that the innovations ✏t are iid rvs with N(μ, 2), μ = 0 andModel II: Xt = 0.6Xt 1 + 0.4Xt 1✏t 1 + ✏t,  2 = 2.0 (a2 + b2 2 = 0.68 < 1),Model III: Xt = 0.6Xt 1 + 0.4Xt 1✏t 1 + ✏t,  2 = 3.0 (a2 + b2 2 = 0.84 < 1),Model IV: Xt = 0.6Xt 1 + 0.7Xt 1✏t 1 + ✏t,  2 = 1.0 (a2 + b2 2 = 0.85 < 1).All these models satisfy the conditions a2 + b2 2 < 1 and |a| < 1, which guarantee stationarity. However, neither of the models satisfy the invertibility condition given in (2.3). Due to this constraint, the CML method implemented before will not be used here. The simulated samples of size n = 5000 are pre- sented in Figure 3 and the statistics calculated with the simulated parameters are presented in Table 3 (N = 106 and kN = 1000). The kernel density esti- mates of a, b and  2 are presented in Figures 4, 5 and 6. For this large sample size, the good agreement between the true values of a, b and  2 and the es- timates obtained clearly show the benefits of using an ABC algorithm in the framework (of the complex) bilinear models. The results obtained for models IIand III indicate that the estimation of a and b is mostly a↵ected by the value of  2.]]></page><page Index="157" isMAC="true"><![CDATA[Model IIIII IVTrue values a = 0.6 b = 0.4  2 = 2.0 a = 0.6 b = 0.4  2 = 3.0 a = 0.6 b = 0.7  2 = 1.0Min Q0.25 Q0.50 Mean 0.3615 0.5325 0.6001 0.59760.1605 0.3362 0.4215 0.4190 0.8264 1.7320 2.1110 2.2110 0.2063 0.4818 0.5573 0.5515 0.1750 0.3277 0.3839 0.3830 1.5976 2.7146 3.2480 3.4051 0.1893 0.4358 0.4899 0.4918 0.2693 0.5505 0.6487 0.6395 0.6535 1.0610 1.2350 1.2980Q0.75 0.66230.4988 2.6380 0.6293 0.4434 3.9411 0.5527 0.7346 1.4940Max0.8271 0.6941 4.2670 0.7909 0.5530 7.1505 0.7923 0.8999 2.59001000.PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 147Density Observations0 1 2 3 4 0 20 40 60Density Observations0.0 0.5 1.0 1.5 2.0 2.5 3.0 −50 0 50 100Density Observations0.0 0.1 0.2 0.3 0.4 0.5 0.6−40 −20 0 20 400 1000 2000 3000 4000 5000 Index0 1000 20003000 4000 50000 1000 2000 3000 IndexFigure 3. Fixed bilinear data sets - Model II (left), Model III (centre) and Model IV (right).Table 3. Statistics for the samples of size kN =0.3 0.4 0.5 0.6 0.7 0.8 0.9 0.2 0.4 0.6 0.8 1N = 1000 Bandwidth = 0.03 N = 1000 Bandwidth = 0.042 34Figure 4. Kernel estimates - model Xt = 0.6Xt 1 + 0.4Xt 1✏t 1 + ✏t, ✏t iid N(0, 2) - a (left), b (centre) and  2 (right) - in all the graphs N represents kN .Index4000 5000s0.0887 0.1083 0.6453 0.105 0.077 0.885 0.0899 0.1277 0.3143N = 1000 Bandwidth = 0.15]]></page><page Index="158" isMAC="true"><![CDATA[148P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMANDensity Density0 1 2 3 4 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5Density Density0.0 0.5 1.0 1.5 2.0 2.5 0 1 2 3 4Density Density0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 0.0 0.1 0.2 0.3 0.40.2 0.4 0.6 0.8 0.1 0.2 0.3 0.4 0.5 0.6 1 2 3 4 5 6 7 8N = 1000 Bandwidth = 0.02369 N = 1000 Bandwidth = 0.02 N = 1000 Bandwidth = 0.2001Figure 5. Kernel density estimates of the accepted values - Xt = 0.6Xt 1 + 0.4Xt 1 ✏t 1 + ✏t , ✏t are iid N (0, 3) - a (left), b (centre) and  2 (right) - in all the graphs N represents kN .0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 1.0 0.5 1.0 1.5 2.0 2.5N = 1000 Bandwidth = 0.01972 N = 1000 Bandwidth = 0.03 N = 1000 Bandwidth = 0.07105Figure 6. Kernel density estimates of the accepted values - Xt = 0.6Xt 1 + 0.7Xt 1 ✏t 1 + ✏t , ✏t are iid N (0, 1) - a (left), b (centre) and  2 (right) - in all the graphs N represents kN .]]></page><page Index="159" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 1495. Comments and conclusionsBilinear processes form a very flexible class models for capturing heavy- tailed, nonlinear features in time series. However, for time series exhibiting heavy tails, the choice of the model for the innovations and the model for the time series, from the general class given in (2.1), is not straightforward. For example, a choice of a Gaussian structure for the innovations would imply a choice of high value of m in the model. Alternatively, the choice of lower order m may simplify the nonlinear structure. However, this comes at a price of choosing a heavy-tailed model for the innovations (Turkman et al. [38]). In either case, least squares methods or conditional likelihood methods for parameter estimation are no longer satisfactory. In this situation, simulation- based methods may work better and we looked at the viability of using an ABC approach with an appropriate set of summary statistics. The method seems to work well for simple first order bilinear processes and is quite promising. However, the bottleneck for extensions to higher order models is the di culty of verifying the invertibility and the stationary conditions, while searching for admissible solutions within the parameter space.References[1] M. A. Beaumont, W. Zhang, and D. J. Balding, Approximate Bayesian computation in population genetics, Genetics 162, 2015–2035, 2002.[2] G. Biau, F. C´erou, and A. Guyader, New insights into Approximate Bayesian Compu- tation, Ann. Inst. Henri Poincar´e Probab. Stat. 51, 376–403, 2015.[3] M. G. B. Blum and O. Franc¸ois, Highly tolerant likelihood-free Bayesian inference: an adaptive non-linear heterocedastic model, Available at arXiv:0810.0896v1, 2008.[4] R. W. Brockett, Non-linear systems and di↵erential geometry. Recent trends in system theory, Proc. IEEE 64 (1), 61–72, 1976.[5] C. W. S. Chen, Bayesian analysis of bilinear time series models: a Gibbs sampling approach, Comm. Statist. Theory Methods 21 (12), 3407–3425, 1992.[6] C. W. S. Chen, Bayesian inferences and forecasting in bilinear time series models, Comm. Statist. Theory Methods 21 (6), 1725–1743, 1992.[7] B. P. Carlin and T. A. Louis, Bayesian Methods for Data Analysis, 3rd Edition, Chap- man & Hall/CRC, Boca Raton, 2009.[8] S. Coles, An Introduction to Statistical Modeling of Extreme Values, Springer-Verlag, London, 2001.[9] K. Csill´ery, M. G. B. Blum, O. E. Gaggiotti, and O. Franc¸ois, Approximate Bayesian Computation (ABC) in practice, Trends Ecol. Evol. 25 (7), 410–418, 2010.[10] A. L. M. Dekkers, J. H. J. Einmahl, and L. de Haan, A moment estimator for the index of an extreme value distribution, Ann. Statist. 17 (4), 1833–1855, 1989.[11] P. Del Moral, A. Doucet, and A. Jasra, Sequential Monte Carlo Samplers, J. R. Stat. Soc. Ser. B Stat. Methodol. 68, 411–436, 2006.[12] A. Doucet, Sequential Importance Sampling Resampling, Lecture 3 at SAMSI, North Carolina, USA, 2010.]]></page><page Index="160" isMAC="true"><![CDATA[150 P. DE ZEA BERMUDEZ, M. A. AMARAL TURKMAN, AND K. F. TURKMAN[13] A. Doucet, N. Freitas, and N. Gordon, An introduction to Sequencial Monte Carlo Methods, in: Sequential Monte Carlo Methods in Practice, A. Doucet, N. Freitas, N. Gordon (eds.), Springer, New York, 3–13, 2001.[14] H. Feng, L. Peng, and F. Zhu, Interval estimation for a simple bilinear model, Statist. Probab. Lett. 83 (10), 2152–2159, 2013.[15] T. Grahn, A conditional least squares approach to bilinear time series estimation, J. Time Ser. Anal. 16 (5), 509–529, 1995.[16] P. J. Green, Reversible jump Markov chain Monte Carlo computation and Bayesian model determination, Biometrika 82, 711–732, 1995.[17] P. Embrechts, C. Klu¨ppelberg, and T. Mikosch, Modelling Extremal Events for Insurance and Finance, Springer-Verlag, Berlin, 1997.[18] P. Marjoram, J. Molitor, V. Plagnol, and S. Tavar´e, Markov chain Monte Carlo without likelihoods, Proc. Natl. Acad. Sci. USA 100 (26), 15324–15328, 2003.[19] M. M. Gabr, Maximum likelihood fitting of bilinear models to time series with missing observations, in: Developments in Time Series Analysis, T. Subba Rao (ed.), Chapman & Hall/CRC, London, 283–291, 1993.[20] M. M. Gabr, Recursive estimation of bilinear time series models, Comm. Statist. Theory Methods 21 (8), 2261—2277, 1992.[21] A. Gelfand and A. Smith, Sampling-bases approaches to calculating marginal densities, J. Amer. Statist. Assoc. 85 (410), 398–409, 1990.[22] S. Geman and D. Geman, Stochastic relaxation, Gibbs distributions, and the bayesian restoration of images, IEEE T. Pattern. Anal. 6, 721—741, 1984.[23] W. R. Gilks, N. G. Best, and K. K. C. Tan, Adaptive rejection Metropolis sampling within Gibbs sampling, J. R. Stat. Soc. Ser. C. Appl. Stat. 44 (4), 455–472, 1995.[24] W. K. Kim, L. Billard, and I. V. Basawa, Estimation for the first-order diagonal bilineartime series model, J. Time Ser. Anal. 11 (3), 215–229, 1990.[25] R. Leon-Gonzalez and F. Yang, Bayesian Inference and Forecasting in the StationaryBilinear Model, University of East Anglia Applied and Financial Economics WorkingPaper Series 055, School of Economics, University of East Anglia, Norwich, UK, 2014.[26] S. Ling, L. Peng, and F. Zhu, Inference for a special bilinear time series model, J. TimeSer. Anal. 36, 61–66, 2014.[27] J. M. Marin, P. Pudlo, C. P. Robert, and R. Ryder, Approximate Bayesian Computa-tional methods, Stat. Comput. 22 (6), 1167–1180, 2012.[28] R. M. Neal, Slice Sampling, Ann. Statist. 31 (3), 705–767, 2003.[29] C. D. Paulino, M. A. Amaral Turkman, and B. Murteira, Estat´ıstica Bayesiana,Funda¸c˜ao Calouste Gulbenkian, Lisboa, 2003.[30] T. D. Pham and L. T. Tran, On the first-order bilinear time series model, J. Appl.Probab. 18 (3), 617–627, 1981.[31] S. Resnick, Modeling Data Networks, in: Extreme Values in Finance, Telecommunica-tions and the Environment, B. Finkensta¨dt, H. Rootz´en (eds.), Chapman & Hall/CRC,Boca Raton, Florida, 287–372, 2003.[32] H. Rue, S. Martino, and N. Chopin, Approximate Bayesian inference for latent Gaussianmodels by using integrated nested Laplace approximations, J. R. Stat. Soc. Ser. B Stat.Methodol. 71 (2), 319–392, 2009.[33] M. G. Scotto, M´etodos de estimac¸a˜o em modelos bilineares, MSc thesis, Faculty ofSciences, Lisbon, 1997.[34] S. A. O. Sesay and T. Subba Rao, Frequency domain estimation of bilinear time seriesmodels, J. Time Ser. Anal. 13 (6), 521–545, 1992.]]></page><page Index="161" isMAC="true"><![CDATA[PARAMETER ESTIMATION OF BILINEAR PROCESSES USING ABC 151[35] T. Subba Rao, On the theory of bilinear time series models, J. Roy. Statist. Soc. Ser. B 43 (2), 244-255, 1981.[36] T. Tubba Rao and M. M. Gabr, An Introduction to Bispectral Analysis and Bilinear Time Series Models, Lecture Notes in Statist. 24, Springer-Verlag, New York, 1984.[37] T. Subba Rao and M. E. Silva, Identification of bilinear time series models BL(p, 0, p, 1), Statist. Sinica 2 (2), 465–478, 1992.[38] K. F. Turkman, M. G. Scotto, and P. de Zea Bermudez, Non-Linear Time Series: Extreme Events and Integer Value Problems, Springer-Verlag, Switzerland, 2014.[39] R. D. Wilkinson, Approximate Bayesian computation (ABC) gives exact results under the assumption of model error, Stat. Appl. Genet. Mol. Biol. 12 (2), 129–141, 2013.[40] K. F. Turkman, and M. A. Amaral Turkman, Extremes of bilinear time series models, J. Time Ser. Anal. 18 (3), 305–320, 1997.(P. de Zea Bermudez, M. A. Amaral Turkman, and K. F. Turkman) DEIO-CEAUL, Faculdade de Cieˆncias, ULisboaE-mail address: pcbermudez@fc.ul.pt; maturkman@fc.ul.pt; kfturkman@fc.ul.pt]]></page><page Index="162" isMAC="true"><![CDATA[]]></page><page Index="163" isMAC="true"><![CDATA[TEXTOS DE MATEMA´TICA1 J.A. GREEN. Classical Invariants, 1993.2 E.M. SA´. Interlacing Problems for Invariant Factors, 1998.3 J.A. GREEN. Classical Groups, 1995.4 J.A. GREEN. Hall Algebras and Quantum Groups, 1994.5 W.TUTSCHKE.InhomogeneousEquationsinComplexAnalysis,1995.6 W. BERGWEILER. An Introduction to Complex Dynamics, 1995.7 J. CNOPS, H. MALONEK. An Introduction to Cli↵ord Analysis, 1995.8 P. MELLON. An Introduction to Several Complex Variables, 1995.9 J.A. GREEN. Shu✏e Algebras, Lie Algebras and Quantum Groups,1995.10 F. CRAVEIRO DE CARVALHO (ed.). Dia de Topologia e Geometria,Coimbra, May 21, 1996.11 F.A. OLIVEIRA, P. OLIVEIRA, M.F. PATR´ICIO, J.A. FERREIRA.(eds.). First Meeting on Numerical Methods for Partial Di↵erentialEquations, Coimbra, 25-27 September 1995, 1997.12 B. BANASCHEWSKI. The Real Numbers in Pointfree Topology, 1997.13 A.P. MOURO, C. VREUGDENHIL. Numerical Methods for Advec-tion-Dominated Problems, 1998.14 J. DE GRAAF. Evolution Equations, 1998.15 J. DE GRAAF. Evolution Equations in Harmonic Function Spaces,1998.16 R. HEERSINK. Initial Value Problems in Scales of Banach Spaces,1998.17 E. WEGERT. Complex Methods for Boundary Value Problems I, 1999.]]></page><page Index="164" isMAC="true"><![CDATA[18 L. VON WOLFERSDORF. Complex Methods for Boundary Value Problems II, 1999.19 A.P. SANTANA, A.L. DUARTE, J.F. QUEIRO´ (eds.). Matrices and Group Representations, Coimbra, 6-8 May 1998, 1999.20 M.J.H. ANTHONISSEN, J. DE GRAAF. Hopper’s Shape Evolution Equation, 1999.21 M. SOBRAL, M.M. CLEMENTINO, J. PICADO, L. SOUSA (eds.). School on Category Theory and Applications, Coimbra, July 13-17, 1999.22 A. DURA´N. Symmetries of Di↵erential Equations and Numerical Ap- plications, 1999.23 I. VAISMAN. Lectures on Symplectic and Poisson Geometry (Notes edited by F. Petalidou), 2000.24 J. VAILLANT, J. CARVALHO E SILVA (eds.). Op´erateurs Di↵´eren- tiels et Physique Math´ematique, 2000.25 P. MELLON. An Introduction to Bounded Symmetric Domains, 2000.26 W. SPRO¨ßIG, K. GU¨RLEBECK. Cli↵ord Analysis and BoundaryValue Problems of Partial Di↵erential Equations, 2000.27 O. MARTIO. Modern Tools in the Theory of Quasiconformal Maps,2000.28 N.J. WILDBERGER. Algebraic Structures Associated to Group Ac-tions and Sums of Hermitian Matrices, 2001.29 E.H. TWIZELL (ed.). Aspects of Computational Bio-Mathematics.(Notes compiled by P. M. Rosa), 2001.30 A.M. ALMEIDA, R. RODRIGUES. Classificac¸˜ao e Complexidade deProblemas, 2001.31 G.THORBERGSSON.GeometryofSubmanifoldsinEuclideanSpaces,2001.]]></page><page Index="165" isMAC="true"><![CDATA[32 H.ALBUQUERQUE,R.CASEIRO,F.J.CRAVEIRODECARVALHO, J.M. NUNES DA COSTA (eds.). The J. A. Pereira da Silva Birthday Schrift, 2002.33 R. W. CARTER. Representations of the Monster, 2002.34 A.J.G. BENTO, A.M. CAETANO, S.D. MOURA, J.S. NEVES (eds.).The J. A. Sampaio Martins Anniversary Volume, 2004.35 F. BORCEAUX. Algumas Teorias de Galois dos Corpos e dos An´eis,2004.36 G. JANK. Matrix Riccati Di↵erential Equations, 2005.37 R.W. CARTER. Cluster Algebras, 2006.38 R. KAHLE, I. OITAVEM (eds.). Days in Logic ’06:Two Tutorials, 2006.39 O. AZENHAS, A.L. DUARTE, J.F. QUEIRO´ , A.P. SANTANA (eds.).Mathematical Papers in Honour of Eduardo Marques de S´a, 2006.40 R. KAHLE. The applicative realm, 2007.41 J. PICADO, A. PULTR. Locales Treated Mostly in a Covariant Way,2008.42 J. BURESˇ, R. LA´VICˇKA, V. SOUCˇEK . Elements of QuaternionicAnalysis and Radon Transform, 2009.43 J. CARDOSO, K. HU¨PER, P. SARAIVA (eds.). Mathematical papersin honour of F´atima Silva Leite, 2011.44 C. FONSECA, S. FURTADO, A. KOVACˇ EC, CHI-KWONG LI (eds.).The Nat´alia Bebiano Anniversary Volume, 2013.45 D. BOURN, N. MARTINS-FERREIRA, A. MONTOLI, M. SOBRAL.Schreier split epimorphisms in monoids and in semirings, 2013.]]></page><page Index="166" isMAC="true"><![CDATA[46 M. M. CLEMENTINO, G. JANELIDZE, J. PICADO, L. SOUSA, W. THOLEN (eds.). Categorical in Algebra and Topology: Special Vol- ume in Honour of Manuela Sobral, 2014.]]></page><page Index="167" isMAC="true"><![CDATA[]]></page><page Index="168" isMAC="true"><![CDATA[]]></page></pages></Search>