Introduction to R

R is a powerful programming language and environment for statistical computing and graphics. It’s widely used in various fields, including linguistics, computer science, and data analysis. This document introduces you to some of the key features and basic concepts of R.

What R Can Do

  • Statistical Analysis: From simple descriptive statistics to complex modeling, R can handle it all.
  • Data Visualization: Create insightful plots and graphs to represent your data visually.
  • Data Manipulation: Clean, transform, and shape your data with ease.
  • Text Analysis: Ideal for linguistics, R can analyze and process textual data.

Basics of R

Basic Operations

First of all, we can use R as a calculator. R follows the standard order of operations (Please Excuse My Dear Aunt Sally! - Parentheses, Exponents, Multiplication and Division, Addition and Subtraction).

Addition
# Assigning values to x and y
x <- 10
y <- 2
# Addition: x + y
addition <- x + y
addition        # 12
## [1] 12

The above code adds x and y together, resulting in 12.

Division
# Division: x / y
division <- x / y
division        # 5
## [1] 5

This divides x by y, resulting in 5.

Modulus
# Modulus (remainder of division): x %% y
modulus <- x %% y
modulus        # 0 
## [1] 0

This calculates the remainder of dividing x by y, resulting in 0.

Integer Division
# Integer division (quotient of division): x %/% y
int_division <- x %/% y
int_division        # 5
## [1] 5

This calculates the integer division of x by y, resulting in 5.

Exponentiation
# Exponentiation: x raised to the power y (x^y)
exponentiation <- x ^ y
exponentiation        # 100
## [1] 100

This raises x to the power of y, resulting in 100.

Variables and Data Types

In R, a variable provides a name for a value, allowing you to store and manipulate data in your code. You can declare variables using the assignment operator (<- or =).

# Assigning values to variables
x <- 10
y = 5
x
## [1] 10
y
## [1] 5

Variables in R can hold different types of data, such as numbers, characters, and logical values. These are known as data types, and understanding them is vital as they govern what you can do with the data.

Variable Naming Rules

When naming variables in R, adhere to the following guidelines: Begin with a letter; numbers and certain punctuation are not allowed at the start.

  • Begin with a letter; numbers and certain punctuation are not allowed at the start.
  • Use only numbers, letters, ., and _; other punctuation is disallowed.
  • Maintain case sensitivity; ‘Variable’ and ‘variable’ are distinct.
  • Aim for descriptive yet concise names (e.g., ‘Speaker’ good; ‘S’ vague; ‘all_my_participant_codes_go_here_plz’ too long).
  • Avoid names that already exist in R (e.g., don’t use ‘pi’). Use exists("name") to check if a name is taken.
  • Do not name a variable “data.”
  • Refer to R4DS2 section 3.3 for more on variable naming styles.

Checking Data Types with class()

The class() function is used to determine the data type of any variable in R. This can be very handy for debugging or understanding how to handle a particular variable.

class(x)  # Output the class of x
## [1] "numeric"
Numeric

Numeric data type is used to store numeric values.

# Example of numeric data type
num <- 42.5
num
## [1] 42.5
class(num)  # Output the class of num
## [1] "numeric"
Integer

Integers are whole numbers without a decimal point.

# Example of integer data type
int <- as.integer(42.5)
int
## [1] 42
class(int)  # Output the class of int
## [1] "integer"
Character

Character data type is used to store strings. It’s important to ensure that strings are enclosed in quotes (either single or double quotes). For example, "Hello, linguistics!" or 'Hello, linguistics!'.

# Example of character data type
char <- "Hello, linguistics!"
char
## [1] "Hello, linguistics!"
class(char)  # Output the class of char
## [1] "character"
Logical/Boolean

Logical data type is used to store TRUE or FALSE values.

# Example of logical data type
log_val <- TRUE
log_val
## [1] TRUE
class(log_val)  # Output the class of log_val
## [1] "logical"
Vector

A vector contains elements of the same type. It’s useful when you want to store and manipulate a collection of similar items.

x <- c(1, 2, 3)
class(x) # "numeric"
## [1] "numeric"
List

A list can contain elements of different types. It’s helpful for grouping related but different types of information together.

my_list <- list(1, "a", TRUE)
class(my_list) # "list"
## [1] "list"
Factor

Factors are used to store categorical data, where the categories are known and limited. This is useful in statistical modeling to represent variables that have a fixed number of different values.

The concept of levels is central to factors. Levels define the possible categories that the factor can take, and they allow for consistent ordering and comparison of categories, even if they are not inherently ordinal.

For example, a factor representing t-shirt sizes could have the levels “small,” “medium,” and “large.” These levels help R understand the relationships between the categories, allowing for meaningful comparisons and analyses.

# Creating a factor with specific levels
sizes <- factor(c("extra-small", "small", "medium", "small", "large", "extra-large"), 
                levels = c("extra-small", "small", "medium", "large", "extra-large"))
class(sizes) # "factor"
## [1] "factor"
levels(sizes) # "extra-small" "small" "medium"  "large" "extra-large"
## [1] "extra-small" "small"       "medium"      "large"       "extra-large"
DataFrame

Data frames are used to store tabular data, where each column can be of a different type. This is particularly useful when dealing with data sets where you have different attributes for each observation, such as spreadsheets or CSV files.

df <- data.frame(name = c("Alice", "Bob"), age = c(25, 30))
class(df) # "data.frame"
## [1] "data.frame"
df
##    name age
## 1 Alice  25
## 2   Bob  30

Loading Packages in R

In R, packages are collections of functions and data sets developed by the community. They enhance the capability of R by adding new functions, methods, and classes. Loading a package in R means that it is loaded into the memory and is available for use.

Why Are Packages Important?

Packages can save time and effort by providing ready-made solutions to common problems. They often include functions that have been optimized and tested, ensuring reliability.

How to Load a Package

You can load a package using the library() function. Before loading, you need to install it using the install.packages() function if it is not already installed.

Installing a Package

(You may just do it in the console)

# Installing the ggplot2 package for data visualization

# install.packages("ggplot2")  ## This line is now necessary for me (it is commented out) because I have already installed this package

Loading a Package

# Loading the ggplot2 package
library(ggplot2)

Upcoming Packages to Explore

tidyverse: An opinionated collection of R packages designed for data science. It includes several packages that make it easier to clean, process, and visualize data.

dplyr: Part of the tidyverse, dplyr provides a set of tools for efficiently manipulating datasets in R.

ggplot2: ggplot2 is a system for declaratively creating graphics and provides a high-level interface for creating attractive and complex plots easily.

These packages will enable us to handle complex data analysis tasks in a more efficient and user-friendly way. As we progress through the weeks, we’ll explore these packages in depth, learning how to apply them to various real-world data scenarios. This will not only enhance your understanding of R but also equip you with the tools needed to perform robust data analysis and visualization.

Loading Data into R

R provides several functions to import data from various file formats. Understanding how to import data is a critical step, as it allows you to bring in external datasets and manipulate them in R.

CSV Files

Comma-Separated Values (CSV) files are a common format for data. You can import CSV files using the read.csv() function:

# Importing a CSV file
data_csv <- read.csv("wk1_Data.csv", # File path
                     header = TRUE,  # Indicates the first row is the header
                     sep = ",",      # Separator used between fields
                     stringsAsFactors = FALSE) # Do not convert strings to factors
head(data_csv) # Display the first few rows, the default is six
##   Sub Gender       L1 TOEFL Gender_raw
## 1   1      1  English    91          M
## 2   2      2 Japanese    93          F
## 3   3      2   Korean    92          f
## 4   4      1   German    94       male
## 5   5      1   German    95       Male
## 6   6      2  English    99     female

SAV Files (SPSS)

SAV files are often used in SPSS. You can import SAV files using the foreign library and the read.spss() function

library(foreign)
data_sav <- read.spss("DeKeyser2000.sav", to.data.frame = TRUE)
head(data_sav)
##   Age GJTScore   Status
## 1   8      170 Under 15
## 2  11      181 Under 15
## 3   9      198 Under 15
## 4  11      194 Under 15
## 5  13      196 Under 15
## 6   4      193 Under 15

Examining Data in R

When working with a dataset in R, it’s important to get a quick overview of its structure and contents. There are several functions available to help you examine your data:

Basic Overview

head() and tail(): These functions allow you to view the first and last observations in your data frame. You can specify the number of observations using the n argument.

# Viewing the first six observations
head(data_csv)
##   Sub Gender       L1 TOEFL Gender_raw
## 1   1      1  English    91          M
## 2   2      2 Japanese    93          F
## 3   3      2   Korean    92          f
## 4   4      1   German    94       male
## 5   5      1   German    95       Male
## 6   6      2  English    99     female
# Viewing the last six observations
tail(data_csv)
##    Sub Gender       L1 TOEFL Gender_raw
## 5    5      1   German    95       Male
## 6    6      2  English    99     female
## 7    7      2 Japanese   100     female
## 8    8      1   Korean    99       male
## 9    9      2  Chinese   100     Female
## 10  10      1  Chinese   101       Male
# Viewing the first and last 10 observations
head(data_csv, n = 10)
##    Sub Gender       L1 TOEFL Gender_raw
## 1    1      1  English    91          M
## 2    2      2 Japanese    93          F
## 3    3      2   Korean    92          f
## 4    4      1   German    94       male
## 5    5      1   German    95       Male
## 6    6      2  English    99     female
## 7    7      2 Japanese   100     female
## 8    8      1   Korean    99       male
## 9    9      2  Chinese   100     Female
## 10  10      1  Chinese   101       Male
tail(data_csv, n = 10)
##    Sub Gender       L1 TOEFL Gender_raw
## 1    1      1  English    91          M
## 2    2      2 Japanese    93          F
## 3    3      2   Korean    92          f
## 4    4      1   German    94       male
## 5    5      1   German    95       Male
## 6    6      2  English    99     female
## 7    7      2 Japanese   100     female
## 8    8      1   Korean    99       male
## 9    9      2  Chinese   100     Female
## 10  10      1  Chinese   101       Male

Dimensions and Names

ncol() and nrow(): These functions provide the number of variables (columns) and observations (rows) in the dataset.

names(): This function retrieves the names of your data frame’s variables.

ncol(data_csv)
## [1] 5
nrow(data_csv)
## [1] 10
names(data_csv)
## [1] "Sub"        "Gender"     "L1"         "TOEFL"      "Gender_raw"

Detailed Information

summary(): Gives a statistical summary of all variables in the data frame.

str(): Checks the structure of the data, including data types.

summary(data_csv)
##       Sub            Gender         L1                TOEFL       
##  Min.   : 1.00   Min.   :1.0   Length:10          Min.   : 91.00  
##  1st Qu.: 3.25   1st Qu.:1.0   Class :character   1st Qu.: 93.25  
##  Median : 5.50   Median :1.5   Mode  :character   Median : 97.00  
##  Mean   : 5.50   Mean   :1.5                      Mean   : 96.40  
##  3rd Qu.: 7.75   3rd Qu.:2.0                      3rd Qu.: 99.75  
##  Max.   :10.00   Max.   :2.0                      Max.   :101.00  
##   Gender_raw       
##  Length:10         
##  Class :character  
##  Mode  :character  
##                    
##                    
## 
str(data_csv)
## 'data.frame':    10 obs. of  5 variables:
##  $ Sub       : int  1 2 3 4 5 6 7 8 9 10
##  $ Gender    : int  1 2 2 1 1 2 2 1 2 1
##  $ L1        : chr  "English" "Japanese" "Korean" "German" ...
##  $ TOEFL     : int  91 93 92 94 95 99 100 99 100 101
##  $ Gender_raw: chr  "M" "F" "f" "male" ...

Modifying Data Types

You can change the data types of specific variables using functions like as.factor() and as.numeric()

data_csv$Sub <- as.factor(data_csv$Sub)    # Change Sub to Factor (nominal)
data_csv$Gender <- as.factor(data_csv$Gender)
data_csv$TOEFL <- as.numeric(data_csv$TOEFL) # Change TOEFL to Numeric (continuous)
str(data_csv) # Double check data type
## 'data.frame':    10 obs. of  5 variables:
##  $ Sub       : Factor w/ 10 levels "1","2","3","4",..: 1 2 3 4 5 6 7 8 9 10
##  $ Gender    : Factor w/ 2 levels "1","2": 1 2 2 1 1 2 2 1 2 1
##  $ L1        : chr  "English" "Japanese" "Korean" "German" ...
##  $ TOEFL     : num  91 93 92 94 95 99 100 99 100 101
##  $ Gender_raw: chr  "M" "F" "f" "male" ...

Summarizing and Tabulating

table(): Helps in counting occurrences of unique values in a variable or creating contingency tables.

unique(): Returns the unique values in a column.

table(data_csv$L1) # Count by L1
## 
##  Chinese  English   German Japanese   Korean 
##        2        2        2        2        2
table(data_csv$Gender) # Count by gender
## 
## 1 2 
## 5 5
table(data_csv$L1, data_csv$Gender) # Contingency table
##           
##            1 2
##   Chinese  1 1
##   English  1 1
##   German   2 0
##   Japanese 0 2
##   Korean   1 1
unique(data_csv$Gender_raw) # Unique values in Gender_raw
## [1] "M"      "F"      "f"      "male"   "Male"   "female" "Female"

Now we can see that gender is represented in various forms like “M”, “male”, “Male”, “F”, “f”, “female”, etc. We need to standardize these values for analysis.

Using the ifelse() Function to Clean Data We can use the ifelse() function to create conditions that transform these different representations into standardized values

grammar rule: ifelse(test, yes, no)

  • test: A logical expression that returns a TRUE or FALSE value.
  • yes: The value to return if the test is TRUE.
  • no: The value to return if the test is FALSE.
data_csv$Gender_clean <- ifelse(data_csv$Gender_raw %in% c("M", "male", "Male"), "Male", "Female")

# This line of code creates a new column 'Gender_clean' in the data_csv data frame.
# It examines the 'Gender_raw' column and checks whether its values are "M", "male", or "Male".
# If any of these values are found, "Male" is assigned to the corresponding element in the 'Gender_clean' column.
# If none of these values are found, "Female" is assigned to the corresponding element in the 'Gender_clean' column.

# Check the cleaned data
unique(data_csv$Gender_clean)
## [1] "Male"   "Female"

The inconsistencies in the Gender_raw column could have been avoided by using controlled input methods when collecting data. For example, providing a dropdown menu or multiple-choice questions instead of allowing participants to manually enter the information can significantly reduce errors and variations.

Subsetting and Exporting Data in R

Subsetting Data

In R, subsetting refers to the process of extracting a part of a dataset. You can subset your data based on specific column numbers or column names. This can be particularly useful when you want to work with only a specific part of a large dataset.

Here are three ways to subset the dataframe data_csv, keeping only the first four columns:

Note that subsetting data in R using the square bracket notation df[rows, cols] allows you to extract specific rows and columns from a data frame

Subsetting based on column numbers

data_csv_new <- data_csv[,c(1,2,3,4,6)]
data_csv_new
##    Sub Gender       L1 TOEFL Gender_clean
## 1    1      1  English    91         Male
## 2    2      2 Japanese    93       Female
## 3    3      2   Korean    92       Female
## 4    4      1   German    94         Male
## 5    5      1   German    95         Male
## 6    6      2  English    99       Female
## 7    7      2 Japanese   100       Female
## 8    8      1   Korean    99         Male
## 9    9      2  Chinese   100       Female
## 10  10      1  Chinese   101         Male

Subsetting using a range

data_csv_new <- data_csv[,c(1:4, 6)]
data_csv_new
##    Sub Gender       L1 TOEFL Gender_clean
## 1    1      1  English    91         Male
## 2    2      2 Japanese    93       Female
## 3    3      2   Korean    92       Female
## 4    4      1   German    94         Male
## 5    5      1   German    95         Male
## 6    6      2  English    99       Female
## 7    7      2 Japanese   100       Female
## 8    8      1   Korean    99         Male
## 9    9      2  Chinese   100       Female
## 10  10      1  Chinese   101         Male

Subsetting based on column names

data_csv_new <- data_csv[, c("Sub","Gender","L1","TOEFL", "Gender_clean")]
data_csv_new
##    Sub Gender       L1 TOEFL Gender_clean
## 1    1      1  English    91         Male
## 2    2      2 Japanese    93       Female
## 3    3      2   Korean    92       Female
## 4    4      1   German    94         Male
## 5    5      1   German    95         Male
## 6    6      2  English    99       Female
## 7    7      2 Japanese   100       Female
## 8    8      1   Korean    99         Male
## 9    9      2  Chinese   100       Female
## 10  10      1  Chinese   101         Male

All three approaches result in a new dataframe data_csv containing only the specified columns.

Exporting Data

Once you’ve subsetted your data and made any necessary changes, you may want to save this new dataset to your computer. You can use the write.csv() function to export the dataframe as a CSV file:

write.csv(data_csv_new, file = "wk1_Data_new.csv", row.names = FALSE)

This code will create a new CSV file named wk1_Data_new.csv, containing the data_csv_new dataframe, without row names. The exported file will be saved in your current working directory, which you can find by running getwd() in R.