By Topic

A local dependence measure and its application to screening for high correlations in large data sets

Sign In

Cookies must be enabled to login.After enabling cookies , please use refresh or reload or ctrl+f5 on the browser for the login options.

Formats Non-Member Member
$31 $13
Learn how you can qualify for the best price for this item!
Become an IEEE Member or Subscribe to
IEEE Xplore for exclusive pricing!
close button

puzzle piece

IEEE membership options for an individual and IEEE Xplore subscriptions for an organization offer the most affordable access to essential journal articles, conference papers, standards, eBooks, and eLearning courses.

Learn more about:

IEEE membership

IEEE Xplore subscriptions

3 Author(s)
Sricharan, K. ; Dept. of EECS, Univ. of Michigan, Ann Arbor, MI, USA ; Hero, A.O. ; Rajaratnam, B.

Correlation screening is frequently the only practical way to discover dependencies in very high dimensional data. In correlation screening a high threshold is applied to the matrix of sample correlation coefficients of the multivariate data. The variables having coefficients that exceed the threshold are called discoveries and are classified to be dependent. The mean number of discoveries and the number of false discoveries in correlation screening problems depend on a information-theoretic measure J, a novel type of information divergence that is a function of the joint density of pairs of variables. It is therefore important to estimate J in order to determine screening thresholds for desired false alarm rates. In this paper, we propose a kernel estimator for J, establish asymptotic consistency and determine the asymptotic distribution of the estimator. These results are used to minimize the MSE of the estimator and to determine confidence intervals on J. We use these results to test for dependence between variables in both simulated data sets and also between email spam harvesters. Finally, we use the estimate of J to determine screening thresholds in correlation screening problems involving gene expression data.

Published in:

Information Fusion (FUSION), 2011 Proceedings of the 14th International Conference on

Date of Conference:

5-8 July 2011