Python and R Programming

R Foundations: Vectors, Matrices, Arrays, Lists and Factors

PGCP-BDA

R object and assignment

An R object holds a typed value and assignment binds a name to that object within an environment.

R vector

An R vector is a one-dimensional homogeneous collection; most R operations apply element by element and preserve vector attributes when rules allow.

R atomic types

R atomic vector types include logical, integer, double, complex, character and raw; coercion chooses a common type when values are combined.

vectorized operation

A vectorized operation applies one declared computation across an array or vector using optimized library loops while following broadcasting.

recycling rule

R repeats shorter vectors in element-wise operations; a nonmultiple length usually signals a likely alignment error.

matrix

A two-dimensional homogeneous R object stored as a vector with row and column dimensions.

array

A homogeneous R object with two or more dimensions recorded in its dim attribute.

list in R

A recursive R vector whose elements may have different types, lengths and nested structures.

factor

An R factor stores categorical observations as integer codes plus an ordered set of level labels.

missing values in R

R represents ordinary missing values with typed NA values and undefined numeric results with NaN

R Vectors and Missing Values

R represents a scalar as a vector of length one. Atomic vector types include logical, integer, double, complex, character and raw. Every element of an atomic vector has one type. Combining unlike values causes coercion toward a type that can hold them. NA represents a missing value, NULL represents absence and NaN is an undefined numeric result.

Arithmetic and comparisons operate element by element. Recycling repeats a shorter vector; a nonmultiple length can hide an error. Indices begin at one. Positive indices select positions, negative indices exclude positions and logical indices select positions whose condition is true. Names permit character indexing. Missing values are located with is.na; comparing a value directly with NA does not produce an ordinary true or false result.

Matrices, Arrays, Lists and Factors

A matrix is an atomic vector with a two-dimensional dim attribute, so all cells share one type. R fills matrices column by column unless byrow=TRUE is specified. Elementwise multiplication uses * while matrix multiplication uses %*%. An array extends the same model to more dimensions. drop=FALSE preserves dimensions during selection.

A list can contain objects of different types and sizes. Single brackets return a sublist while double brackets extract one element. $ extracts a named component. A factor represents categories with integer codes and a set of levels. Ordered factors add level order. Converting a numeric factor directly with as.numeric returns its codes rather than numeric labels. Filtering may leave unused levels that can be dropped explicitly.

Vectorized Computation and Attributes

Many R functions accept complete vectors and return complete vectors, avoiding explicit element loops. Vectorization describes an interface and often uses optimized internal operations. Logical conditions can select and replace several positions at once. ifelse is vectorized while if expects one condition. Summary functions such as sum, mean, min and max return missing results when input contains NA unless missing values are removed deliberately.

Objects carry attributes. Names label vector elements. Dimensions turn a vector into a matrix or array. Class affects method dispatch and printing. Changing attributes does not necessarily copy or validate the underlying values in the same way as reconstructing an object. str reveals internal structure and is often more informative than printed output.

Matrix indexing with one index follows the underlying column-major vector. Two indices identify row and column. Row and column names provide semantic access. Matrix operations require conformable dimensions. Transpose uses t, solving a linear system uses solve and binding functions combine compatible rows or columns.

Lists support recursive structures. Data frames are lists of equal-length columns with class and row-name attributes. Model objects are often named lists containing coefficients, residuals and metadata. Understanding the difference between [, [[ and $ is essential because one returns a container while the others normally return its component.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.