Python and R Programming
R Functions, Statistics, Visualization and R Markdown
PGCP-BDA
R function
An R function captures its defining lexical environment, binds actual arguments lazily to formal parameters and returns the value of its last evaluated.
lexical scope in R
R resolves a free symbol through the chain of environments where a function was defined, rather than the call stack from which it happened to be invoked.
descriptive statistics in R
R functions that summarize center, spread, frequency, quantiles and missingness without fitting a predictive model.
statistical model in R
A fitted R object produced from a formula and data, containing estimated parameters, diagnostics and prediction methods.
base graphics
R’s built-in stateful plotting system in which a plot is opened and then modified with additional drawing functions.
ggplot2 grammar
A layered plotting model that maps variables to aesthetics and combines geometric marks, statistics, scales, coordinates and themes.
R Markdown
A document format combining Markdown prose, executable code chunks and rendered output in a reproducible report.
reproducible report
A report whose narrative, calculations and figures are regenerated from the same code, data and declared environment.
R and Python comparison
R emphasizes vectorized statistics and formula-based analysis, while Python provides a broad general-purpose programming ecosystem.
Ggplot Grammar
A grammar of graphics composes data, aesthetic mappings, geometric marks, statistical transformations, scales, coordinates and facets. A concise Python expression is useful only when its data flow remains readable. Choose the built-in type or library abstraction that matches ordering, uniqueness, lookup, numerical or tabular requirements. Observe return values and side effects and keep transformation code separate from input, storage and presentation. When using ggplot grammar, document which object owns mutable state, which caller releases resources and which failures can propagate. Names and types should express the contract without forcing a reader to inspect every implementation detail. Automated tests should verify the public behavior and avoid depending on incidental internal ordering unless that ordering is part of the contract.
Test Statistic
A test statistic reduces sample evidence to a quantity whose distribution under the null hypothesis is known or approximated. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat test statistic as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Visual Encoding
A visual encoding maps data values to position, length, angle, color, size or shape; position and aligned length usually support more accurate comparison. Python’s runtime model matters: names refer to objects, operations are dispatched by type and mutability determines whether an operation changes an object or creates another one. Clear code preserves object boundaries, validates external input and uses exceptions to report conditions a caller can handle. A small example of visual encoding should be traced from object or input creation through every relevant operation and return value. Include an empty or null-like case and one invalid case so the exception or boundary behavior is visible. Production use should keep external input validation, business logic, storage and presentation in separate functions or classes.
Labels Scales And Legends
Titles, units, tick scales, annotations and legends must identify what is measured and prevent visual comparison from implying false magnitude. Python’s runtime model matters: names refer to objects, operations are dispatched by type and mutability determines whether an operation changes an object or creates another one. Clear code preserves object boundaries, validates external input and uses exceptions to report conditions a caller can handle. A small example of labels scales and legends should be traced from object or input creation through every relevant operation and return value. Include an empty or null-like case and one invalid case so the exception or boundary behavior is visible. Production use should keep external input validation, business logic, storage and presentation in separate functions or classes.
Robust Statistic
A robust statistic changes relatively little when a small fraction of observations are extreme or contaminated. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For robust statistic, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Standard Deviation
Standard deviation is the square root of variance and expresses typical spread in the original measurement units. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat standard deviation as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Skewed Data
Skewed data has an asymmetric distribution, often making the mean and standard deviation less representative than robust summaries. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat skewed data as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Quartile And Percentile
A percentile is a value below which a stated proportion falls and quartiles are the 25th, 50th and 75th percentiles. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat quartile and percentile as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Type I And Type Ii Error
A type I error rejects a true null; a type II error fails to reject a false null and their trade-off depends on effect size, noise and sample size. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat type I and type II error as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Skewness
Skewness summarizes distribution asymmetry, but sample estimates can be unstable and should be interpreted with plots and domain context. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat skewness as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Pearson Correlation
Pearson correlation standardizes covariance to the range minus one to one and measures linear association rather than causation. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat Pearson correlation as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
P-Value
A p-value is the probability, assuming the null model and test procedure, of obtaining a statistic at least as incompatible with the null as the observed statistic. A useful implementation separates definition, computation and interpretation. First state what the quantity or mechanism means; then apply the correct procedure; finally ask whether the result supports the original question. Sensitivity checks reveal whether small changes in data, parameters or assumptions produce an unstable conclusion. In an end-to-end workflow, p-value belongs between documented input checks and an explicit interpretation step. Save parameters, transformations, random seeds and evaluation splits so another person can reproduce the result. Operational use also needs thresholds, exception handling, monitoring and a response when incoming data no longer resembles the validated conditions.
Outlier
An outlier is an observation unusually distant under a chosen representation; it may be an error, a rare valid case or evidence of another process. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat outlier as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Data Investigation
Data investigation traces suspicious values to their source, definition, collection process and downstream effect before deciding whether to alter them. The concept becomes practical when its inputs are measured consistently and its result is compared with a meaningful baseline. Training performance or a visually attractive pattern is insufficient evidence. Validation must reflect the intended population, preserve time and group boundaries where relevant and report limitations that affect action. A common error is to treat data investigation as an automatic conclusion rather than a model of limited evidence. Separate observations used to construct the result from observations used to test it. Compare with a simple reference, examine important subgroups and report what the method cannot determine. This keeps mathematical correctness connected to decision quality.
Variance
Variance is the average squared deviation from a mean under the chosen population or sample convention and measures squared spread. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For variance, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Null And Alternative Hypothesis
The null hypothesis defines a reference claim or model, while the alternative represents departures the test is designed to detect. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For null and alternative hypothesis, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Mean Median And Mode
The mean balances all values, the median divides ordered observations and the mode is the most frequent value; each describes a different center. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For mean median and mode, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Range And Interquartile Range
Range is maximum minus minimum, while interquartile range spans the middle half and is less sensitive to extreme observations. A useful implementation separates definition, computation and interpretation. First state what the quantity or mechanism means; then apply the correct procedure; finally ask whether the result supports the original question. Sensitivity checks reveal whether small changes in data, parameters or assumptions produce an unstable conclusion. In an end-to-end workflow, range and interquartile range belongs between documented input checks and an explicit interpretation step. Save parameters, transformations, random seeds and evaluation splits so another person can reproduce the result. Operational use also needs thresholds, exception handling, monitoring and a response when incoming data no longer resembles the validated conditions.
Influence
Influence measures how much an estimate or fitted model changes when an observation is perturbed or removed. A useful implementation separates definition, computation and interpretation. First state what the quantity or mechanism means; then apply the correct procedure; finally ask whether the result supports the original question. Sensitivity checks reveal whether small changes in data, parameters or assumptions produce an unstable conclusion. In an end-to-end workflow, influence belongs between documented input checks and an explicit interpretation step. Save parameters, transformations, random seeds and evaluation splits so another person can reproduce the result. Operational use also needs thresholds, exception handling, monitoring and a response when incoming data no longer resembles the validated conditions.
Coefficient Of Variation
The coefficient of variation divides standard deviation by a nonzero mean to compare relative spread on ratio scales. A useful implementation separates definition, computation and interpretation. First state what the quantity or mechanism means; then apply the correct procedure; finally ask whether the result supports the original question. Sensitivity checks reveal whether small changes in data, parameters or assumptions produce an unstable conclusion. In an end-to-end workflow, coefficient of variation belongs between documented input checks and an explicit interpretation step. Save parameters, transformations, random seeds and evaluation splits so another person can reproduce the result. Operational use also needs thresholds, exception handling, monitoring and a response when incoming data no longer resembles the validated conditions.
Spurious Correlation
A spurious correlation appears meaningful but is produced by chance, a common trend, multiple testing, confounding or selection. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For spurious correlation, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Significance Level
The significance level is a preselected bound on the probability of rejecting a true null under repeated valid use of the test. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For significance level, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Robust Treatment
Robust treatment may correct confirmed errors, transform variables, cap values with justification, use robust methods or analyze rare cases separately. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For robust treatment, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Covariance
Covariance is the average product of centered variables and indicates the direction of linear co-variation, with magnitude dependent on units. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For covariance, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Chi-Square Test
A chi-square test compares observed and expected counts for goodness of fit or independence when count and sampling assumptions are satisfied. To use the idea correctly, identify the population or state being described, the information available at decision time and the quantity being optimized or estimated. Check assumptions against the data instead of treating a mathematical name as proof. A result should be interpreted with its uncertainty, failure modes and the cost of a wrong decision. For chi-square test, a worked analysis should state the unit being studied, how every input was obtained and what comparison would show that the method adds value. Numerical output must retain units and uncertainty. If the population, measurement process or operating conditions change, the earlier conclusion may no longer hold and should be validated again.
Z-Test
A z-test compares a standardized estimate with the standard normal distribution when its variance assumptions or large-sample approximation are justified. A useful implementation separates definition, computation and interpretation. First state what the quantity or mechanism means; then apply the correct procedure; finally ask whether the result supports the original question. Sensitivity checks reveal whether small changes in data, parameters or assumptions produce an unstable conclusion. In an end-to-end workflow, z-test belongs between documented input checks and an explicit interpretation step. Save parameters, transformations, random seeds and evaluation splits so another person can reproduce the result. Operational use also needs thresholds, exception handling, monitoring and a response when incoming data no longer resembles the validated conditions.
Correlation And Causation
Correlation records association, while causal claims require a defensible design that addresses confounding, selection, time order and alternative explanations. A useful implementation separates definition, computation and interpretation. First state what the quantity or mechanism means; then apply the correct procedure; finally ask whether the result supports the original question. Sensitivity checks reveal whether small changes in data, parameters or assumptions produce an unstable conclusion. In an end-to-end workflow, correlation and causation belongs between documented input checks and an explicit interpretation step. Save parameters, transformations, random seeds and evaluation splits so another person can reproduce the result. Operational use also needs thresholds, exception handling, monitoring and a response when incoming data no longer resembles the validated conditions.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.