Skip to content
Coming soon: Ally. Your guide to the world of AI and R. Learn More →
R for the Rest of Us Logo

Getting Started With R

Import Data Again

Transcript

Click on the transcript to go to that point in the video. Please note that transcripts are auto generated and may contain minor inaccuracies.

View code shown in video
library(tidyverse)

coffee_ratings <- read_csv(
    "coffee_ratings.csv",
    na = c("No rating", "Unknown", "8", "1", "3", "4", "23", "47"),
    col_types = cols(total_cup_points = "d", harvest_year = "i")
)

coffee_ratings

Your Turn

  1. Adjust your read_csv() code so that you import the data again

  2. Use the na argument to tell read_csv() what should be treated as missing in the altitude_mean_meterscolumn

  3. Use the col_types argument to make sure that altitude_mean_meters gets imported as numeric data

  4. Examine your data again to confirm that the changes were made successfully

The code that I created and that you can use to get started is below.

library(tidyverse)

coffee_ratings <- read_csv(
    "coffee_ratings.csv",
    na = c("No rating", "Unknown", "8", "1", "3", "4", "23", "47"),
    col_types = cols(total_cup_points = "d", harvest_year = "i")
)

coffee_ratings

Have any questions? Put them below and we will help you out!

You need to be signed-in to comment on this post. Login.

Caitlin McLemore

Caitlin McLemore • March 3, 2026

What is the difference between the data types of number, double, and integer?

Gracielle Higino

Gracielle Higino Coach • March 6, 2026

Hi Caitlin! Both double and integer are types of numbers. Integers (sometimes you can see the notations Int, In16, Int 32 or Int64) refer to whole numbers, without decimals, such as 123, and can represent variables such as years, age and - to be very niche here - sample plots in space. They can often be thought of as categorical variables. Doubles are numbers with decimal digits (like 123.5), often thought of as continuous variables, and can represent variables such as height, length, and weight.

Eda Akpek

Eda Akpek • March 13, 2026

I am trying to do the practice exercise for altitude_mean_meters. The code shows that it's executing, but I don't see any changes in the histogram.

Gracielle Higino

Gracielle Higino Coach • March 14, 2026

Hi Eda! Do you want to share the exact code you used? The change in the histogram is really subtle, but you should see differences when opening the little "toggle" for more info if you define what should be read as NAs in the dataset. But you're right that changing the column type doesn't change much in this case, but I suggest that you play a bit with the arguments to see what happens in each case. What happens if you don't designate the NA values for that column? What happens if you do? What if you change the column type?

Eda Akpek

Eda Akpek • March 16, 2026

I figured it out! I set the altitude_mean_meters as an integer, and then, because it changed -999.00 to -999, I was able to use -999 in my code to set as "NA."

Gracielle Higino

Gracielle Higino Coach • March 16, 2026

Awesome! 🎉

Al-Afroza Sultana

Al-Afroza Sultana • March 26, 2026

Hi, as this is a small data set, we could check the data manually in an ascending or descending manner to find out if there's any abnormality. But how do we check or find out any abnormal values in a comparatively large dataset using R? Thanks.

Gracielle Higino

Gracielle Higino Coach • March 26, 2026

There are a few methods to do that! Here's an informative tutorial to inspect outliers with basic min and max functions and with plots: https://statsandr.com/blog/outliers-detection-in-r/

Plotting is one of the most effective ways to detect outliers and discrepancies - very useful in my area, where we deal with millions of biodiversity data points!

One other tool you can take a look at is OpenRefine, which can be integrated with your R scripts: https://openrefine.org/ OpenRefine is great to detect typos.

Lalitha Vaishnavi Subramanyan

Lalitha Vaishnavi Subramanyan • May 1, 2026

Inputting 1, 4, 23 from harvest_year etc into na worked here, but with a dataset where those can be legitimate values in another column this approach may fail and corrupt data, is there another approach where we would avoid this pitfall?

Gracielle Higino

Gracielle Higino Coach • May 7, 2026

Great questions! Yes, in this case you could skip the na assignment when reading the data, and deal with it in a case-by-case level, perhaps, using mutate(), or case_when(), or na_if(). See some examples here: https://dplyr.tidyverse.org/reference/na_if.html#ref-examples

Kathy Worley

Kathy Worley • September 24, 2026

I keep getting this error and nothing runs any suggestions Error in original_source():

Gracielle Higino

Gracielle Higino Coach • September 23, 2026

Hi Kathy! Could you share the whole code you used and the complete error message that you've got? You can also run rlang::last_trace() right after you get the error and copy-paste the message here so we have more clues about what's causing this error.

Kathy Worley

Kathy Worley • September 24, 2026

library(tidyverse)

coffee_ratings <- read_csv("coffee_ratings.csv", na=c("no rating","unknown","8","1","3","4","23","47","-999", col_types=cols(total_cup_points)="d", harvest_year="I")) glimpse(coffee_ratings) library(skimr) skim (coffee_ratings) library(pointblank) scan_data (coffee_ratings) Error in original_source(): ! c:/Users/kathyw/OneDrive - Conservancy of Southwest Florida/Desktop/import.r:5:33: unexpected '=' 4: na=c("no rating","unknown","8","1","3","4","23","47","-999", 5: col_types=cols(total_cup_points)=

Gracielle Higino

Gracielle Higino Coach • September 24, 2026

YAY we've found it! Your code has a "=" in the wrong place at line 5.

You've got:

col_types=cols(total_cup_points)="d",
harvest_year="I"))

But you should have: col_types = cols(total_cup_points = "d", harvest_year = "i")

You're closing a parenthesis after total_cup_points, which messes up a bit with the syntax of the function and makes you add an extra parenthesis after harvest_year = "I". Also, i should be lower case.

Let me know if this helps!

Kathy Worley

Kathy Worley • September 27, 2026

still getting errors. I think it is something with the input file I am going to try removing and reinstalling everything and see if that helps

Kathy Worley

Kathy Worley • September 27, 2026

still getting errors. I think it is something with the input file I am going to try removing and reinstalling everything and see if that helps

Gracielle Higino

Gracielle Higino Coach • September 28, 2026

Ok! If you still get errors, feel free to copy-paste the error message and the exact code you used here. [=

Maia Werner-Avidon

Maia Werner-Avidon • September 24, 2026

I am stuck. Trying to get rid of the outliers in the altitude column by adding to David's example code, but it does not remover the outliers. coffee_ratings <- read_csv( "coffee_ratings.csv", na = c("No rating", "Unknown", "8", "1", "3", "4", "23", "47", "-999.00", "190164.00", "110000.00"), col_types = cols(total_cup_points = "d", harvest_year = "i") )

Gracielle Higino

Gracielle Higino Coach • September 24, 2026

Hi Maia! That's a very subtle problem: the thing is that R is trying to find the exact value "110000.00", however, although it's displayed like that in the data viewer, this number does not exist in the dataset. What exists is "110000", without the digits. You find this out when you copy-paste from the data viewer. Try this out: try copying the "110000.00" from the altitude_mean_meters column and paste it on the console or on your script. You'll see that the .00 will not paste over. Now try copying the first row value of the total_cup_points. It comes with the extra digits, right?

However, if you add these numbers to the NA list without the quotation marks, you'll get the result you want, either with or without the digits. That's because in R, 110000.00 without quotes is a numeric literal, and R stores it as the number 110000. Test this:

1.00 == 1

"1.00" == 1

You'll see R doesn't treat these values as the same thing.

In the process of reading your CSV, the na argument in read_csv() matches the raw text in the file, before any parsing happens. The function reads each cell as a string, checks whether that string is exactly one of your na values, and only then converts the rest to double, integer, and so on. That's why "11000.00" won't match what you see as 11000.00.

Long story short, you just need to remove the extra digits from your outliers (making R match the exact string value on your dataset) or remove the quotation marks around them (making R read them as a number).

I know it can be confusing! Let me know if it's not clear.

Sarah Aparicio

Sarah Aparicio • September 28, 2026

For harvest_year, is there a way to mark all non 4 digit numbers as missing? Instead of manually listing out 8, 1, 3, 4... etc.

Gracielle Higino

Gracielle Higino Coach • September 28, 2026

Hi Sarah! Technically, yes, but you'll need more advanced functions, some of them we'll learn on Week 2. The best approach would be to transform these values into NAs after you read the dataset, as a transformation of your raw data. Also, NA assignments should be treated carefully, so unless you are absolutely sure you will never need numbers that are smaller, and that the values are entered correctly (in some cases, 0042 can be read as a 4-digits number, but you'd like it removed), you would probably be better off just assigning specific values for your NAs. Ideally you'd have this designed into your data collection step!