5
매우 많은 수의 데이터 포인트에서 값을 대치하는 방법은 무엇입니까?
데이터 세트가 매우 커서 약 5 %의 임의 값이 없습니다. 이 변수들은 서로 상관되어 있습니다. 다음 예제 R 데이터 세트는 더미 상관 데이터가있는 장난감 예제 일뿐입니다. set.seed(123) # matrix of X variable xmat <- matrix(sample(-1:1, 2000000, replace = TRUE), ncol = 10000) colnames(xmat) <- paste ("M", 1:10000, sep ="") rownames(xmat) …
12
r
random-forest
missing-data
data-imputation
multiple-imputation
large-data
definition
moving-window
self-study
categorical-data
econometrics
standard-error
regression-coefficients
normal-distribution
pdf
lognormal
regression
python
scikit-learn
interpolation
r
self-study
poisson-distribution
chi-squared
matlab
matrix
r
modeling
multinomial
mlogit
choice
monte-carlo
indicator-function
r
aic
garch
likelihood
r
regression
repeated-measures
simulation
multilevel-analysis
chi-squared
expected-value
multinomial
yates-correction
classification
regression
self-study
repeated-measures
references
residuals
confidence-interval
bootstrap
normality-assumption
resampling
entropy
cauchy
clustering
k-means
r
clustering
categorical-data
continuous-data
r
hypothesis-testing
nonparametric
probability
bayesian
pdf
distributions
exponential
repeated-measures
random-effects-model
non-independent
regression
error
regression-to-the-mean
correlation
group-differences
post-hoc
neural-networks
r
time-series
t-test
p-value
normalization
probability
moments
mgf
time-series
model
seasonality
r
anova
generalized-linear-model
proportion
percentage
nonparametric
ranks
weighted-regression
variogram
classification
neural-networks
fuzzy
variance
dimensionality-reduction
confidence-interval
proportion
z-test
r
self-study
pdf