Python and R Programming
R Data Frames, Packages, Data Import and Tidy Transformation
PGCP-BDA
R data frame
An R data frame is a tabular list of equal-length vectors whose columns may have different atomic types.
R package and library
An R package is an installable unit of code and data; a library is the package directory or the act of loading a package.
data import in R
Reading external data with an appropriate parser, types, missing-value rules, locale and encoding.
indexing in R
R indexing selects by positive positions, negative exclusions, logical masks, names or matrix coordinates.
apply family
R’s apply-family functions express repeated computation over margins, list elements or groups and differ in input simplification and returned structure.
dplyr verbs
dplyr uses select, filter, mutate, arrange, summarize and group_by to express column-oriented table transformations with consistent data-frame semantics.
tidyr reshaping
tidyr reshapes tables with operations such as pivot_longer and pivot_wider while keeping each variable.
joins in R
Operations that combine data frames by matching key columns, including inner, left, right and full joins.
grouped summarization
Partitioning rows by key and computing one or more aggregate values for each group.
R data quality
Assessment and correction of missingness, invalid types, duplicates, inconsistent categories and out-of-range values in R.
Data Frames and Packages
A data frame is rectangular: its columns have equal length but may have different types. Each row should represent one observation and each column one variable in tidy data. Indexing may use positions, names or logical conditions. Missing values require explicit treatment with is.na and options such as na.rm=TRUE.
An R package contains functions, documentation, metadata and possibly compiled code or datasets. install.packages installs it while library attaches it for a session. A namespace controls exports and imports. package::function identifies a function without attaching the whole package and resolves name conflicts. Reproducible analysis records package versions and the R version.
Import and Tidy Transformation
Delimited import requires a delimiter, header rule, quoting convention, encoding and missing-value representation. Type inference must be checked because identifiers with leading zeroes, dates and mixed columns are easily misclassified. Verify row count, names, types, key uniqueness and missingness after import.
Filtering selects rows, selecting chooses columns, mutating derives columns, arranging orders rows and summarising reduces groups. Grouping changes the unit of aggregation. Joins require declared keys and understood cardinality because duplicate keys multiply rows. Wide-to-long transformation gathers repeated measurement columns into variable-value pairs while long-to-wide transformation spreads key values into columns. Reliable transformation states both input and output grain.
Grouped Data and Joins
Grouped transformation retains grouping metadata until explicitly removed or replaced. A grouped summarise creates one or more rows per group according to the expressions. A grouped mutation returns the original row count while computing within each group. Unexpected grouping can therefore change later results even when the visible columns look unchanged.
An inner join keeps matching keys. A left join keeps every left row. Full and right joins preserve different unmatched sets. Semi joins select left rows that have a match without adding right columns. Anti joins select those without a match. Before joining, keys should be checked for uniqueness on the side expected to contain one row per key. Many-to-many matches produce the Cartesian set within each key.
Dates and times require a defined parser, time zone and treatment of ambiguous local times. Text normalization must not erase meaningful distinctions in identifiers. Numeric parsing should identify currency symbols, thousands separators and decimal conventions rather than silently producing missing values.
An analysis script should preserve raw input, write derived data to a separate location and record parameters. Random operations need a controlled seed where repeatability matters. Exported files should retain column order, names, encoding and missing-value conventions expected by downstream consumers.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.