Import Data Again
This lesson is called Import Data Again, part of the Getting Started With R course. This lesson is called Import Data Again, part of the Getting Started With R course.
Transcript
Click on the transcript to go to that point in the video. Please note that transcripts are auto generated and may contain minor inaccuracies.
Loading transcript...
View code shown in video
library(tidyverse)
coffee_ratings <- read_csv(
"coffee_ratings.csv",
na = c("No rating", "Unknown", "8", "1", "3", "4", "23", "47"),
col_types = cols(total_cup_points = "d", harvest_year = "i")
)
coffee_ratings
Your Turn
Adjust your
read_csv()code so that you import the data againUse the
naargument to tellread_csv()what should be treated as missing in thealtitude_mean_meterscolumnUse the
col_typesargument to make sure thataltitude_mean_metersgets imported as numeric dataExamine your data again to confirm that the changes were made successfully
The code that I created and that you can use to get started is below.
library(tidyverse)
coffee_ratings <- read_csv(
"coffee_ratings.csv",
na = c("No rating", "Unknown", "8", "1", "3", "4", "23", "47"),
col_types = cols(total_cup_points = "d", harvest_year = "i")
)
coffee_ratings
You need to be signed-in to comment on this post. Login.
Caitlin McLemore • March 3, 2026
What is the difference between the data types of number, double, and integer?
Gracielle Higino Coach • March 6, 2026
Hi Caitlin! Both double and integer are types of numbers. Integers (sometimes you can see the notations Int, In16, Int 32 or Int64) refer to whole numbers, without decimals, such as 123, and can represent variables such as years, age and - to be very niche here - sample plots in space. They can often be thought of as categorical variables. Doubles are numbers with decimal digits (like 123.5), often thought of as continuous variables, and can represent variables such as height, length, and weight.
Eda Akpek • March 13, 2026
I am trying to do the practice exercise for altitude_mean_meters. The code shows that it's executing, but I don't see any changes in the histogram.
Gracielle Higino Coach • March 14, 2026
Hi Eda! Do you want to share the exact code you used? The change in the histogram is really subtle, but you should see differences when opening the little "toggle" for more info if you define what should be read as NAs in the dataset. But you're right that changing the column type doesn't change much in this case, but I suggest that you play a bit with the arguments to see what happens in each case. What happens if you don't designate the NA values for that column? What happens if you do? What if you change the column type?
Eda Akpek • March 16, 2026
I figured it out! I set the altitude_mean_meters as an integer, and then, because it changed -999.00 to -999, I was able to use -999 in my code to set as "NA."
Gracielle Higino Coach • March 16, 2026
Awesome! 🎉
Al-Afroza Sultana • March 26, 2026
Hi, as this is a small data set, we could check the data manually in an ascending or descending manner to find out if there's any abnormality. But how do we check or find out any abnormal values in a comparatively large dataset using R? Thanks.
Gracielle Higino Coach • March 26, 2026
There are a few methods to do that! Here's an informative tutorial to inspect outliers with basic min and max functions and with plots: https://statsandr.com/blog/outliers-detection-in-r/
Plotting is one of the most effective ways to detect outliers and discrepancies - very useful in my area, where we deal with millions of biodiversity data points!
One other tool you can take a look at is OpenRefine, which can be integrated with your R scripts: https://openrefine.org/ OpenRefine is great to detect typos.
Lalitha Vaishnavi Subramanyan • May 1, 2026
Inputting 1, 4, 23 from harvest_year etc into na worked here, but with a dataset where those can be legitimate values in another column this approach may fail and corrupt data, is there another approach where we would avoid this pitfall?
Gracielle Higino Coach • May 7, 2026
Great questions! Yes, in this case you could skip the na assignment when reading the data, and deal with it in a case-by-case level, perhaps, using mutate(), or case_when(), or na_if(). See some examples here: https://dplyr.tidyverse.org/reference/na_if.html#ref-examples
Kathy Worley • September 24, 2026
I keep getting this error and nothing runs any suggestions Error in
original_source():Gracielle Higino Coach • September 23, 2026
Hi Kathy! Could you share the whole code you used and the complete error message that you've got? You can also run
rlang::last_trace()right after you get the error and copy-paste the message here so we have more clues about what's causing this error.Kathy Worley • September 24, 2026
library(tidyverse)
coffee_ratings <- read_csv("coffee_ratings.csv", na=c("no rating","unknown","8","1","3","4","23","47","-999", col_types=cols(total_cup_points)="d", harvest_year="I")) glimpse(coffee_ratings) library(skimr) skim (coffee_ratings) library(pointblank) scan_data (coffee_ratings) Error in
original_source(): ! c:/Users/kathyw/OneDrive - Conservancy of Southwest Florida/Desktop/import.r:5:33: unexpected '=' 4: na=c("no rating","unknown","8","1","3","4","23","47","-999", 5: col_types=cols(total_cup_points)=Gracielle Higino Coach • September 24, 2026
YAY we've found it! Your code has a "=" in the wrong place at line 5.
You've got:
But you should have:
col_types = cols(total_cup_points = "d", harvest_year = "i")You're closing a parenthesis after
total_cup_points, which messes up a bit with the syntax of the function and makes you add an extra parenthesis afterharvest_year = "I". Also,ishould be lower case.Let me know if this helps!
Kathy Worley • September 27, 2026
still getting errors. I think it is something with the input file I am going to try removing and reinstalling everything and see if that helps
Kathy Worley • September 27, 2026
still getting errors. I think it is something with the input file I am going to try removing and reinstalling everything and see if that helps
Gracielle Higino Coach • September 28, 2026
Ok! If you still get errors, feel free to copy-paste the error message and the exact code you used here. [=
Maia Werner-Avidon • September 24, 2026
I am stuck. Trying to get rid of the outliers in the altitude column by adding to David's example code, but it does not remover the outliers. coffee_ratings <- read_csv( "coffee_ratings.csv", na = c("No rating", "Unknown", "8", "1", "3", "4", "23", "47", "-999.00", "190164.00", "110000.00"), col_types = cols(total_cup_points = "d", harvest_year = "i") )
Gracielle Higino Coach • September 24, 2026
Hi Maia! That's a very subtle problem: the thing is that R is trying to find the exact value "110000.00", however, although it's displayed like that in the data viewer, this number does not exist in the dataset. What exists is "110000", without the digits. You find this out when you copy-paste from the data viewer. Try this out: try copying the "110000.00" from the
altitude_mean_meterscolumn and paste it on the console or on your script. You'll see that the.00will not paste over. Now try copying the first row value of thetotal_cup_points. It comes with the extra digits, right?However, if you add these numbers to the NA list without the quotation marks, you'll get the result you want, either with or without the digits. That's because in R, 110000.00 without quotes is a numeric literal, and R stores it as the number 110000. Test this:
You'll see R doesn't treat these values as the same thing.
In the process of reading your CSV, the
naargument inread_csv()matches the raw text in the file, before any parsing happens. The function reads each cell as a string, checks whether that string is exactly one of your na values, and only then converts the rest to double, integer, and so on. That's why "11000.00" won't match what you see as 11000.00.Long story short, you just need to remove the extra digits from your outliers (making R match the exact string value on your dataset) or remove the quotation marks around them (making R read them as a number).
I know it can be confusing! Let me know if it's not clear.
Sarah Aparicio • September 28, 2026
For harvest_year, is there a way to mark all non 4 digit numbers as missing? Instead of manually listing out 8, 1, 3, 4... etc.
Gracielle Higino Coach • September 28, 2026
Hi Sarah! Technically, yes, but you'll need more advanced functions, some of them we'll learn on Week 2. The best approach would be to transform these values into NAs after you read the dataset, as a transformation of your raw data. Also, NA assignments should be treated carefully, so unless you are absolutely sure you will never need numbers that are smaller, and that the values are entered correctly (in some cases, 0042 can be read as a 4-digits number, but you'd like it removed), you would probably be better off just assigning specific values for your NAs. Ideally you'd have this designed into your data collection step!