R is a powerful programming language and environment for statistical computing and graphics. It’s widely used in various fields, including linguistics, computer science, and data analysis. This document introduces you to some of the key features and basic concepts of R.
First of all, we can use R as a calculator. R follows the standard order of operations (Please Excuse My Dear Aunt Sally! - Parentheses, Exponents, Multiplication and Division, Addition and Subtraction).
# Assigning values to x and y
x <- 10
y <- 2
# Addition: x + y
addition <- x + y
addition # 12
## [1] 12
The above code adds x and y together, resulting in 12.
# Division: x / y
division <- x / y
division # 5
## [1] 5
This divides x by y, resulting in 5.
# Modulus (remainder of division): x %% y
modulus <- x %% y
modulus # 0
## [1] 0
This calculates the remainder of dividing x by
y, resulting in 0.
# Integer division (quotient of division): x %/% y
int_division <- x %/% y
int_division # 5
## [1] 5
This calculates the integer division of x by
y, resulting in 5.
# Exponentiation: x raised to the power y (x^y)
exponentiation <- x ^ y
exponentiation # 100
## [1] 100
This raises x to the power of y, resulting
in 100.
In R, a variable provides a name for a value, allowing you to store
and manipulate data in your code. You can declare variables using the
assignment operator (<- or =).
# Assigning values to variables
x <- 10
y = 5
x
## [1] 10
y
## [1] 5
Variables in R can hold different types of data, such as numbers, characters, and logical values. These are known as data types, and understanding them is vital as they govern what you can do with the data.
When naming variables in R, adhere to the following guidelines: Begin with a letter; numbers and certain punctuation are not allowed at the start.
., and _; other
punctuation is disallowed.exists("name") to check if a name is taken.Checking Data Types with class()
The class() function is used to determine the data type
of any variable in R. This can be very handy for debugging or
understanding how to handle a particular variable.
class(x) # Output the class of x
## [1] "numeric"
Numeric data type is used to store numeric values.
# Example of numeric data type
num <- 42.5
num
## [1] 42.5
class(num) # Output the class of num
## [1] "numeric"
Integers are whole numbers without a decimal point.
# Example of integer data type
int <- as.integer(42.5)
int
## [1] 42
class(int) # Output the class of int
## [1] "integer"
Character data type is used to store strings. It’s important to
ensure that strings are enclosed in quotes (either single or double
quotes). For example, "Hello, linguistics!" or
'Hello, linguistics!'.
# Example of character data type
char <- "Hello, linguistics!"
char
## [1] "Hello, linguistics!"
class(char) # Output the class of char
## [1] "character"
Logical data type is used to store TRUE or FALSE values.
# Example of logical data type
log_val <- TRUE
log_val
## [1] TRUE
class(log_val) # Output the class of log_val
## [1] "logical"
A vector contains elements of the same type. It’s useful when you want to store and manipulate a collection of similar items.
x <- c(1, 2, 3)
class(x) # "numeric"
## [1] "numeric"
A list can contain elements of different types. It’s helpful for grouping related but different types of information together.
my_list <- list(1, "a", TRUE)
class(my_list) # "list"
## [1] "list"
Factors are used to store categorical data, where the categories are known and limited. This is useful in statistical modeling to represent variables that have a fixed number of different values.
The concept of levels is central to factors. Levels define the possible categories that the factor can take, and they allow for consistent ordering and comparison of categories, even if they are not inherently ordinal.
For example, a factor representing t-shirt sizes could have the levels “small,” “medium,” and “large.” These levels help R understand the relationships between the categories, allowing for meaningful comparisons and analyses.
# Creating a factor with specific levels
sizes <- factor(c("extra-small", "small", "medium", "small", "large", "extra-large"),
levels = c("extra-small", "small", "medium", "large", "extra-large"))
class(sizes) # "factor"
## [1] "factor"
levels(sizes) # "extra-small" "small" "medium" "large" "extra-large"
## [1] "extra-small" "small" "medium" "large" "extra-large"
Data frames are used to store tabular data, where each column can be of a different type. This is particularly useful when dealing with data sets where you have different attributes for each observation, such as spreadsheets or CSV files.
df <- data.frame(name = c("Alice", "Bob"), age = c(25, 30))
class(df) # "data.frame"
## [1] "data.frame"
df
## name age
## 1 Alice 25
## 2 Bob 30
In R, packages are collections of functions and data sets developed by the community. They enhance the capability of R by adding new functions, methods, and classes. Loading a package in R means that it is loaded into the memory and is available for use.
Packages can save time and effort by providing ready-made solutions to common problems. They often include functions that have been optimized and tested, ensuring reliability.
You can load a package using the library() function.
Before loading, you need to install it using the
install.packages() function if it is not already
installed.
(You may just do it in the console)
# Installing the ggplot2 package for data visualization
# install.packages("ggplot2") ## This line is now necessary for me (it is commented out) because I have already installed this package
# Loading the ggplot2 package
library(ggplot2)
tidyverse: An opinionated collection of R packages designed for data science. It includes several packages that make it easier to clean, process, and visualize data.
dplyr: Part of the tidyverse, dplyr provides a set of tools for efficiently manipulating datasets in R.
ggplot2: ggplot2 is a system for declaratively creating graphics and provides a high-level interface for creating attractive and complex plots easily.
These packages will enable us to handle complex data analysis tasks in a more efficient and user-friendly way. As we progress through the weeks, we’ll explore these packages in depth, learning how to apply them to various real-world data scenarios. This will not only enhance your understanding of R but also equip you with the tools needed to perform robust data analysis and visualization.
R provides several functions to import data from various file formats. Understanding how to import data is a critical step, as it allows you to bring in external datasets and manipulate them in R.
Comma-Separated Values (CSV) files are a common format for data. You
can import CSV files using the read.csv() function:
# Importing a CSV file
data_csv <- read.csv("wk1_Data.csv", # File path
header = TRUE, # Indicates the first row is the header
sep = ",", # Separator used between fields
stringsAsFactors = FALSE) # Do not convert strings to factors
head(data_csv) # Display the first few rows, the default is six
## Sub Gender L1 TOEFL Gender_raw
## 1 1 1 English 91 M
## 2 2 2 Japanese 93 F
## 3 3 2 Korean 92 f
## 4 4 1 German 94 male
## 5 5 1 German 95 Male
## 6 6 2 English 99 female
SAV files are often used in SPSS. You can import SAV files using the
foreign library and the read.spss() function
library(foreign)
data_sav <- read.spss("DeKeyser2000.sav", to.data.frame = TRUE)
head(data_sav)
## Age GJTScore Status
## 1 8 170 Under 15
## 2 11 181 Under 15
## 3 9 198 Under 15
## 4 11 194 Under 15
## 5 13 196 Under 15
## 6 4 193 Under 15
When working with a dataset in R, it’s important to get a quick overview of its structure and contents. There are several functions available to help you examine your data:
head() and tail(): These functions allow you to view
the first and last observations in your data frame. You can specify the
number of observations using the n argument.
# Viewing the first six observations
head(data_csv)
## Sub Gender L1 TOEFL Gender_raw
## 1 1 1 English 91 M
## 2 2 2 Japanese 93 F
## 3 3 2 Korean 92 f
## 4 4 1 German 94 male
## 5 5 1 German 95 Male
## 6 6 2 English 99 female
# Viewing the last six observations
tail(data_csv)
## Sub Gender L1 TOEFL Gender_raw
## 5 5 1 German 95 Male
## 6 6 2 English 99 female
## 7 7 2 Japanese 100 female
## 8 8 1 Korean 99 male
## 9 9 2 Chinese 100 Female
## 10 10 1 Chinese 101 Male
# Viewing the first and last 10 observations
head(data_csv, n = 10)
## Sub Gender L1 TOEFL Gender_raw
## 1 1 1 English 91 M
## 2 2 2 Japanese 93 F
## 3 3 2 Korean 92 f
## 4 4 1 German 94 male
## 5 5 1 German 95 Male
## 6 6 2 English 99 female
## 7 7 2 Japanese 100 female
## 8 8 1 Korean 99 male
## 9 9 2 Chinese 100 Female
## 10 10 1 Chinese 101 Male
tail(data_csv, n = 10)
## Sub Gender L1 TOEFL Gender_raw
## 1 1 1 English 91 M
## 2 2 2 Japanese 93 F
## 3 3 2 Korean 92 f
## 4 4 1 German 94 male
## 5 5 1 German 95 Male
## 6 6 2 English 99 female
## 7 7 2 Japanese 100 female
## 8 8 1 Korean 99 male
## 9 9 2 Chinese 100 Female
## 10 10 1 Chinese 101 Male
ncol() and nrow(): These functions provide the number of variables (columns) and observations (rows) in the dataset.
names(): This function retrieves the names of your data frame’s variables.
ncol(data_csv)
## [1] 5
nrow(data_csv)
## [1] 10
names(data_csv)
## [1] "Sub" "Gender" "L1" "TOEFL" "Gender_raw"
summary(): Gives a statistical summary of all variables in the data frame.
str(): Checks the structure of the data, including data types.
summary(data_csv)
## Sub Gender L1 TOEFL
## Min. : 1.00 Min. :1.0 Length:10 Min. : 91.00
## 1st Qu.: 3.25 1st Qu.:1.0 Class :character 1st Qu.: 93.25
## Median : 5.50 Median :1.5 Mode :character Median : 97.00
## Mean : 5.50 Mean :1.5 Mean : 96.40
## 3rd Qu.: 7.75 3rd Qu.:2.0 3rd Qu.: 99.75
## Max. :10.00 Max. :2.0 Max. :101.00
## Gender_raw
## Length:10
## Class :character
## Mode :character
##
##
##
str(data_csv)
## 'data.frame': 10 obs. of 5 variables:
## $ Sub : int 1 2 3 4 5 6 7 8 9 10
## $ Gender : int 1 2 2 1 1 2 2 1 2 1
## $ L1 : chr "English" "Japanese" "Korean" "German" ...
## $ TOEFL : int 91 93 92 94 95 99 100 99 100 101
## $ Gender_raw: chr "M" "F" "f" "male" ...
You can change the data types of specific variables using functions
like as.factor() and as.numeric()
data_csv$Sub <- as.factor(data_csv$Sub) # Change Sub to Factor (nominal)
data_csv$Gender <- as.factor(data_csv$Gender)
data_csv$TOEFL <- as.numeric(data_csv$TOEFL) # Change TOEFL to Numeric (continuous)
str(data_csv) # Double check data type
## 'data.frame': 10 obs. of 5 variables:
## $ Sub : Factor w/ 10 levels "1","2","3","4",..: 1 2 3 4 5 6 7 8 9 10
## $ Gender : Factor w/ 2 levels "1","2": 1 2 2 1 1 2 2 1 2 1
## $ L1 : chr "English" "Japanese" "Korean" "German" ...
## $ TOEFL : num 91 93 92 94 95 99 100 99 100 101
## $ Gender_raw: chr "M" "F" "f" "male" ...
table(): Helps in counting occurrences of unique values in a variable or creating contingency tables.
unique(): Returns the unique values in a column.
table(data_csv$L1) # Count by L1
##
## Chinese English German Japanese Korean
## 2 2 2 2 2
table(data_csv$Gender) # Count by gender
##
## 1 2
## 5 5
table(data_csv$L1, data_csv$Gender) # Contingency table
##
## 1 2
## Chinese 1 1
## English 1 1
## German 2 0
## Japanese 0 2
## Korean 1 1
unique(data_csv$Gender_raw) # Unique values in Gender_raw
## [1] "M" "F" "f" "male" "Male" "female" "Female"
Now we can see that gender is represented in various forms like “M”, “male”, “Male”, “F”, “f”, “female”, etc. We need to standardize these values for analysis.
Using the ifelse() Function to Clean Data We can use the
ifelse() function to create conditions that transform these
different representations into standardized values
grammar rule: ifelse(test, yes, no)
data_csv$Gender_clean <- ifelse(data_csv$Gender_raw %in% c("M", "male", "Male"), "Male", "Female")
# This line of code creates a new column 'Gender_clean' in the data_csv data frame.
# It examines the 'Gender_raw' column and checks whether its values are "M", "male", or "Male".
# If any of these values are found, "Male" is assigned to the corresponding element in the 'Gender_clean' column.
# If none of these values are found, "Female" is assigned to the corresponding element in the 'Gender_clean' column.
# Check the cleaned data
unique(data_csv$Gender_clean)
## [1] "Male" "Female"
The inconsistencies in the Gender_raw column could have
been avoided by using controlled input methods when collecting data. For
example, providing a dropdown menu or multiple-choice questions instead
of allowing participants to manually enter the information can
significantly reduce errors and variations.
In R, subsetting refers to the process of extracting a part of a dataset. You can subset your data based on specific column numbers or column names. This can be particularly useful when you want to work with only a specific part of a large dataset.
Here are three ways to subset the dataframe data_csv,
keeping only the first four columns:
Note that subsetting data in R using the square bracket notation df[rows, cols] allows you to extract specific rows and columns from a data frame
Subsetting based on column numbers
data_csv_new <- data_csv[,c(1,2,3,4,6)]
data_csv_new
## Sub Gender L1 TOEFL Gender_clean
## 1 1 1 English 91 Male
## 2 2 2 Japanese 93 Female
## 3 3 2 Korean 92 Female
## 4 4 1 German 94 Male
## 5 5 1 German 95 Male
## 6 6 2 English 99 Female
## 7 7 2 Japanese 100 Female
## 8 8 1 Korean 99 Male
## 9 9 2 Chinese 100 Female
## 10 10 1 Chinese 101 Male
Subsetting using a range
data_csv_new <- data_csv[,c(1:4, 6)]
data_csv_new
## Sub Gender L1 TOEFL Gender_clean
## 1 1 1 English 91 Male
## 2 2 2 Japanese 93 Female
## 3 3 2 Korean 92 Female
## 4 4 1 German 94 Male
## 5 5 1 German 95 Male
## 6 6 2 English 99 Female
## 7 7 2 Japanese 100 Female
## 8 8 1 Korean 99 Male
## 9 9 2 Chinese 100 Female
## 10 10 1 Chinese 101 Male
Subsetting based on column names
data_csv_new <- data_csv[, c("Sub","Gender","L1","TOEFL", "Gender_clean")]
data_csv_new
## Sub Gender L1 TOEFL Gender_clean
## 1 1 1 English 91 Male
## 2 2 2 Japanese 93 Female
## 3 3 2 Korean 92 Female
## 4 4 1 German 94 Male
## 5 5 1 German 95 Male
## 6 6 2 English 99 Female
## 7 7 2 Japanese 100 Female
## 8 8 1 Korean 99 Male
## 9 9 2 Chinese 100 Female
## 10 10 1 Chinese 101 Male
All three approaches result in a new dataframe data_csv containing only the specified columns.
Once you’ve subsetted your data and made any necessary changes, you
may want to save this new dataset to your computer. You can use the
write.csv() function to export the dataframe as a CSV
file:
write.csv(data_csv_new, file = "wk1_Data_new.csv", row.names = FALSE)
This code will create a new CSV file named
wk1_Data_new.csv, containing the
data_csv_new dataframe, without row names. The exported
file will be saved in your current working directory, which you can find
by running getwd() in R.