Membuka halaman seterusnya.
Memaparkan bab seterusnya…
Psst… sesuaikan cara membaca Anda.
Fon dan tema tersedia di Tampilan. Keselesaan mata anda juga penting.
Memaparkan bab seterusnya…
How two prices move together. In stress, many alts correlate near 1 with Bitcoin.
Menyemak sokongan bacaan suara pada pelayar ini…
Bacaan ini kini tersedia dalam bahasa Inggeris. Antara muka menggunakan bahasa pilihan anda.
Baca teks asal bahasa Inggeris →In statistics, correlation is a type of statistical relationship between two random variables or bivariate data. It usually refers to the extent to which a pair of quantities are linearly related. More generally, an arbitrary relationship between variables is called an association, meaning the degree to which the variability in one can be accounted for by the other.
The presence of a correlation is not sufficient to infer the presence of a causal relationship, and this is often stated as "correlation does not imply causation". Furthermore, the concept of correlation is not the same as dependence: if two variables are independent, then they are uncorrelated, but the opposite is not necessarily true – even if two variables are uncorrelated, they might be dependent on each other.
Correlations are useful because they can indicate a predictive relationship that can be exploited in practice. For example, an electrical utility may produce less power on a mild day based on the correlation between electricity demand and weather. In this example, there is a causal relationship, because extreme weather causes people to use more electricity for heating or cooling.
The most familiar measure of dependence between two quantities is the Pearson product-moment correlation coefficient, most commonly called 'Pearson's correlation coefficient' or simply 'the correlation coefficient' (as it is the most common variant). It is obtained by taking the ratio of the covariance between two variables of a numerical dataset normalized to the square root of their variances. Equivalently, Pearson's correlation coefficient can be calculated by dividing the covariance of the two variables by the product of their standard deviations.
Karl Pearson developed the coefficient from a similar idea by Francis Galton.
A Pearson product-moment correlation coefficient attempts to establish a line of best fit through a dataset of two variables by essentially laying out the expected values and the resulting Pearson's correlation coefficient indicates how far away the actual dataset is from the expected values. Depending on the sign of the Pearson's correlation coefficient, the result can be either a negative or positive correlation if there is any sort of relationship between the variables in the data set.
If the variables are independent, Pearson's correlation coefficient is 0. However, because the correlation coefficient detects only linear dependencies between two variables, the converse is not necessarily true. A correlation coefficient of 0 does not imply that the variables are independent.
Even though uncorrelated data do not necessarily imply independence, one can check if random variables are independent if their mutual information is 0.
Rank correlation coefficients, such as Spearman's rank correlation coefficient and Kendall's rank correlation coefficient (τ) measure the extent to which, as one variable increases, the other variable tends to increase, without requiring that increase to be represented by a linear relationship. If, as the one variable increases, the other decreases, the rank correlation coefficients will be negative.
It is common to regard these rank correlation coefficients as alternatives to Pearson's coefficient, used either to reduce the amount of calculation or to make the coefficient less sensitive to non-normality in distributions. However, this view has little mathematical basis, as rank correlation coefficients measure a different type of relationship than the Pearson product-moment correlation coefficient, and are best seen as measures of a different type of association, rather than as an alternative measure of the population correlation coefficient.
The conventional dictum that "correlation does not imply causation" means that correlation cannot be used by itself to infer a causal relationship between the variables. This dictum should not be taken to mean that correlations cannot indicate the potential existence of causal relations. However, the causes underlying the correlation, if any, may be indirect and unknown, and high correlations also overlap with identity relations (tautologies), where no causal process exists (e.g., between two variables measuring the same construct).
Consequently, a correlation between two variables is not a sufficient condition to establish a causal relationship (in either direction).
A correlation between age and height in children is fairly causally transparent, but a correlation between mood and health in people is less so. Does improved mood lead to improved health, or does good health lead to good mood, or both? Or does some other factor underlie both? In other words, a correlation can be taken as evidence for a possible causal relationship, but cannot indicate what the causal relationship, if any, might be.
Dipilih dan diformat ulang daripada Correlation, oleh para kontributornya, dengan lesen CC BY-SA 4.0. Semakan 1376167080. Bahagian dan format telah diringkas; semakan berpaut menyediakan konteks lengkap dan sejarah penyumbang. Teks rujukan ini tetap menggunakan lesen yang sama. Pautan rujukan tambahannya diimport daripada semakan tersebut dan belum disemak secara bebas di sini.