This web log is my online notebook for software tips and hacks to make my life and hopefully that of someone else just a little bit easy. Should you find any mistakes or have any useful suggestions, please do not hesitate to write to me. Peace! Pavan Ghatty
Showing posts with label R. Show all posts
Showing posts with label R. Show all posts
Thursday, January 16, 2014
Wednesday, November 6, 2013
C style formatting in R
Numbering/naming files in a controllable format is important when making animations of a series of plots (by using animate for example). When the plots are being generated in R, there is an easy way to control the names of the output files as shown below;
The advantage here is that when you feed in files with the same base.name and numbered 000..100 (instead of 0..100), the command convert automatically reads them in increasing order w/o you relying on time stamps and such.
The advantage here is that when you feed in files with the same base.name and numbered 000..100 (instead of 0..100), the command convert automatically reads them in increasing order w/o you relying on time stamps and such.
Friday, August 30, 2013
Subsetting data with string matching in R
Getting a subset from a larger pool of data is a major strength in R. This subsetting can be based on numbers or strings; the latter being more challenging. String matching (using grep, grepl, sub) and subsetting (using, well, subset in R) are two separate feature and one can combine them to great effect. Details of what grep, grepl etc can do are at
http://stat.ethz.ch/R-manual/R-devel/library/base/html/grep.html . This site
http://atomicules.co.uk/2010/10/09/invert-a-subset-selection-when-using-grepl-in-r.html
was of great help when I wanted (had) to use grepl in subset but I also wanted to use the invert feature available only in grep and not in grepl. Turns out grepl returns a vectors of T and F which can be easily inverted with a simple !grepl(...)
http://stat.ethz.ch/R-manual/R-devel/library/base/html/grep.html . This site
http://atomicules.co.uk/2010/10/09/invert-a-subset-selection-when-using-grepl-in-r.html
was of great help when I wanted (had) to use grepl in subset but I also wanted to use the invert feature available only in grep and not in grepl. Turns out grepl returns a vectors of T and F which can be easily inverted with a simple !grepl(...)
Thursday, August 15, 2013
Thursday, August 8, 2013
Tuesday, April 23, 2013
Curve fitting in R
This is a simple example of a linear fit where a perfect fit is distorted with the help of a gaussian function. This data is then fit to a linear model. The most important factor determining the success of a fit is the model equation and the second most important factor is suitable starting values of the fit variables (a and b in this case). Further, if you assign the output to a variable called fit, values a and b can be extracted using:
fit < - nls
coef(fit)[1]
coef(fit)[2]
Friday, April 19, 2013
Saving data frames in R
If 'df' is your data frame,
save(df,file="yourfile.rdf",ascii=TRUE)
save the data frame in to a .rdf file which is readable. Although the file as lot more data than you might want, loading the file back into R irons out any wrinkles.
load("yourfile.rdf")
http://stackoverflow.com/questions/8345759/how-to-save-a-data-frame-in-r
Edit Jan 2014
If however you really want to avoid additional lines in the saved file, try the following:
write.table(df,filename,row.names=FALSE,col.names=FALSE)
This should produce a clean file with your data in a neat column.
save(df,file="yourfile.rdf",ascii=TRUE)
save the data frame in to a .rdf file which is readable. Although the file as lot more data than you might want, loading the file back into R irons out any wrinkles.
load("yourfile.rdf")
http://stackoverflow.com/questions/8345759/how-to-save-a-data-frame-in-r
Edit Jan 2014
If however you really want to avoid additional lines in the saved file, try the following:
write.table(df,filename,row.names=FALSE,col.names=FALSE)
This should produce a clean file with your data in a neat column.
Thursday, January 31, 2013
Sweave, R and Latex
If you use R for your statistical analysis and graph plotting needs and then use those figures in your Latex documents, you will find this interesting. Sweave is a package in R which lets takes in an .rnw file (which contains usual text from a tex document with R commands embedded in it), compiles it (in R) and gives out a tex file. This tex file can in turn be compiled by pdflatex (if you have figures) or plain latex to get a properly formatted pdf/ps/dvi file with relevant R output. There are many excellent examples on the interwebs and I strongly recommend using Sweave right away. The beauty of this approach is that on can save a lot of time by not spending on choosing the appropriate dimensions of a a png or an eps output. Sweave does it for you.
Edit: If you want to insert a R-generated figure then add this to your tex file:
\begin{figure}
\begin{center}
<
\end{center}
\end{figure}
If you want to insert text from R but properly tabbed and with all the ampersands added, then add these lines to your tex file:
<
Edit: If you want to insert a R-generated figure then add this to your tex file:
\begin{figure}
\begin{center}
<
\end{center}
\end{figure}
If you want to insert text from R but properly tabbed and with all the ampersands added, then add these lines to your tex file:
<
Tuesday, October 9, 2012
Controlling axis properties in R
http://www.programmingr.com/content/controlling-margins-and-axes-oma-and-mgp/
The mgp option
In addition to changing the margin size of your charts, you may also want to change the way axes and labels are spatially arranged. One method of doing so is the
The mgp option
In addition to changing the margin size of your charts, you may also want to change the way axes and labels are spatially arranged. One method of doing so is the
mgp parameter option. The mgp
setting is defined by a three item vector wherein the first value
represents the distance of the axis labels or titles from the axes, the
second value is the distance of the tick mark labels from the axes, and
the third is the distance of the tick mark symbols from the axes. As
with the oma option discussed above, the distances are given in line widths. The defaults for the mgp setting are c(3, 1, 0). The examples below illustrate the effects of changing the various mgp values. Note: the mgp.axis() function in the Hmisc package can be used to change these settings for each axis individually.Thursday, August 23, 2012
Memory convservation in R
Depending on how large a data set you are handling, R can take up large swaths of your RAM and slow down other processes. It is therefore a good idea to remove unused variables.
ls() # Gives the list of variables/arrays currently in use.
rm(list=ls()) # Removes all the variables and frees up memory.
Useful links:
http://www.matthewckeller.com/html/memory.html
ls() # Gives the list of variables/arrays currently in use.
rm(list=ls()) # Removes all the variables and frees up memory.
Useful links:
http://www.matthewckeller.com/html/memory.html
Thursday, August 2, 2012
Reading/processing multiple files in R
j<-1;
for (i in seq(300,390,15)) {
df<-read.table(paste("sasa",i,".txt",sep=``));
# Files are named "sasa300.txt" for ex.
tab[j,]<-cbind(i,mean(df[,2]),sd(df[,2]));
print(tab[j,]);
j<-j+1;
}
> tab
[,1] [,2] [,3]
[1,] 300 3408.514 96.48097
[2,] 315 3422.388 119.14861
[3,] 330 3391.590 81.57252
[4,] 345 3424.176 84.84211
[5,] 360 3394.205 100.37193
[6,] 375 3462.469 122.60486
[7,] 390 3438.328 92.70519
for (i in seq(300,390,15)) {
df<-read.table(paste("sasa",i,".txt",sep=``));
# Files are named "sasa300.txt" for ex.
tab[j,]<-cbind(i,mean(df[,2]),sd(df[,2]));
print(tab[j,]);
j<-j+1;
}
> tab
[,1] [,2] [,3]
[1,] 300 3408.514 96.48097
[2,] 315 3422.388 119.14861
[3,] 330 3391.590 81.57252
[4,] 345 3424.176 84.84211
[5,] 360 3394.205 100.37193
[6,] 375 3462.469 122.60486
[7,] 390 3438.328 92.70519
Wednesday, July 25, 2012
Quick and dirty way to plot error bars in R
If you have data that looks as follows, a one line command can draw error bars to the plot.
-x---T-----y-------------error-----
p5 315 2.513402 0.5711628
p5 330 2.173868 0.4711594
p5 345 2.132954 0.634045
p5 360 2.092686 0.4851832
p5 375 2.966971 0.9165328
p5 390 2.214886 0.4288291
r<-read.table("data.txt")
plot(r[,2],r[,3],ylim=c(0,4))
lines(r[,2],r[,3])
for(i in 1:length(r[,1])) {segments(r[i,2],r[i,3]+r[i,4],r[i,2],r[i,3]-r[i,4])}
I am drawing a simple line from (x,y+err) to (x,y-err). The result is not as fancy as other efforts but it is enough for my purposes.
-x---T-----y-------------error-----
p5 315 2.513402 0.5711628
p5 330 2.173868 0.4711594
p5 345 2.132954 0.634045
p5 360 2.092686 0.4851832
p5 375 2.966971 0.9165328
p5 390 2.214886 0.4288291
r<-read.table("data.txt")
plot(r[,2],r[,3],ylim=c(0,4))
lines(r[,2],r[,3])
for(i in 1:length(r[,1])) {segments(r[i,2],r[i,3]+r[i,4],r[i,2],r[i,3]-r[i,4])}
I am drawing a simple line from (x,y+err) to (x,y-err). The result is not as fancy as other efforts but it is enough for my purposes.
Wednesday, July 4, 2012
Empty plots in R
When plotting a grid of graphs in R, I needed leave some cells empty. For this the following commands come handy;
plot(0,xaxt='n',yaxt='n',bty='n',pch='',ylab='',xlab='')
plot.new()
grid.newpage()
grid()
Wednesday, June 20, 2012
Subsetting in R
Following are some example of subsetting in R I have used recently.
> head(crystal)
V1 V2 V3 V4
1 1ACJ PHE120 290.980 289.101
2 1ACJ PHE153 307.673 283.814
3 1ACJ PHE155 296.153 89.152
4 1ACJ PHE186 311.606 292.186
5 1ACJ PHE187 298.920 270.613
6 1ACJ PHE197 66.678 323.351
>subset(crystal, crystal$V2=="PHE330" & grepl("^2A",crystal$V1))
http://www.ats.ucla.edu/stat/r/faq/subset_R.htm
http://stackoverflow.com/questions/2125231/subsetting-in-r-using-or-condition-with-strings
> head(crystal)
V1 V2 V3 V4
1 1ACJ PHE120 290.980 289.101
2 1ACJ PHE153 307.673 283.814
3 1ACJ PHE155 296.153 89.152
4 1ACJ PHE186 311.606 292.186
5 1ACJ PHE187 298.920 270.613
6 1ACJ PHE197 66.678 323.351
>subset(crystal, crystal$V2=="PHE330" & grepl("^2A",crystal$V1))
http://www.ats.ucla.edu/stat/r/faq/subset_R.htm
http://stackoverflow.com/questions/2125231/subsetting-in-r-using-or-condition-with-strings
Friday, February 10, 2012
Getting header names in R and playing with csv files
names(myTable)
a<-read.csv("file.csv", header=T, sep=",")
names(a)
motility<-subset(a, Putative.orf.function=="Motility & Attachment")
write.csv(motolity, file="motility.csv")
Edit Feb 15th 2012
grep("Motility", a$Putative.orf.function)
gives the list of rows whose Putative.orf.function column contains the string "Motility".
Using this, one can get the entire row information.
df<-read.csv("file.csv", header=T, sep=",")
names(df)
> names(df)
[1] "Strain.Name" "ID" "Location"
[4] "PA.ORF" "Gene" "Gene.Abbrev."
[7] "Putative.orf.function" "Position.in.ORF" "Frame"
[10] "Genome.Position" "Transposon.Direction" "Transposon"
[13] "F.Primer.Name" "F.Primer.Seq" "F.Primer.Position"
[16] "F.Primer.Tm" "Insert...F" "R.Primer.Name"
[19] "R.Primer.Seq" "R.Primer.Position" "R.Primer.Tm"
[22] "R...Insert" "WT.PCR.Len" "Confirmed."
[25] "How.confirmed." "Notes"
> df$Strain.Name[grep("Attachment",df$Putative.orf.function)]
[1] PW10299 PW10300 PW1728 PW1729 PW1730 PW1731 PW1749 PW1750 PW1751
[10] PW1752 PW1753 PW1754 PW1755 PW1756 PW1757 PW2791 PW2792 PW2794
[19] PW2795 PW2941 PW2942 PW2943 PW2944 PW2945 PW2946 PW2947 PW2948
...
[118] PW8681 PW8682 PW9465 PW9466 PW9467 PW9468 PW9469 PW9470 PW9471 [127] PW9472 PW9473 PW9474
Tuesday, December 6, 2011
Min-Max on R plots
d <- density(rnorm(1000))
plot(d)
d$x[which.max(d$y)]
abline(v=d$x[which.max(d$y)])
http://r.789695.n4.nabble.com/Find-x-value-of-density-plots-td4165926.html
plot(d)
d$x[which.max(d$y)]
abline(v=d$x[which.max(d$y)])
http://r.789695.n4.nabble.com/Find-x-value-of-density-plots-td4165926.html
Monday, November 28, 2011
Thursday, August 18, 2011
png figures in a loop in R
This sample script creates 100 png figures with the iteration number included in the name.
for(i in 1:100) {
a<-rnorm(100)
name<-paste("test",i,".png",sep="")
png(name)
plot(a/max(a),type="l")
dev.off()
}
Friday, February 12, 2010
Plot layout in R
I recently needed to make a figure with two plots side-by-side. Plot 1 was to have a width 3 times that of plot 2. The layout feature in R seemed to be the answer. But, due to either lack of decent tutorials online or just plain bad luck in finding one, I found it rather difficult to understand how this layout function works. The usual help command help(layout) wasn't terribly helpful either. Here go the set of commands that really nailed it for me. Run them on your terminal to follow the rest of the blog.
nf <- layout(matrix(c(1,2),1,2),widths=c(1,1)); layout.show(nf)
Pretty straight forward isn't it? Now try,
nf <- layout(matrix(c(1,2),1,2),widths=c(3,1)); layout.show(nf)
Basically the matrix(c(1,2),1,2) is what contains the layout information. c(1,2) is simply the label of the plot in question. The 1,2 after that tells you the number if rows(1) and columns(2) in the layout. Now a little more complicated example,
nf <- layout(matrix(c(1,3,4,2,5,6,7,8),2,4),widths=c(1,1,1,5)); layout.show(nf)
Hope this helps.
Thursday, February 11, 2010
Colors in R
Hello World,
R is a pretty wicked statistical analysis tool with great resources for plotting data. I use it on a regular basis and I'll use this blog to document notable finds along the way. This will hopefully help me find answers to previously encountered problems by taking a peek at the blog rather than fishing through my sea of bookmarks. Here is a short comment on colors in R.
colors() gives a list of all the available colors in R. It looks like this:
[1] "white" "aliceblue" "antiquewhite"
[4] "antiquewhite1" "antiquewhite2" "antiquewhite3"
[7] "antiquewhite4" "aquamarine" "aquamarine1"
[10] "aquamarine2" "aquamarine3" "aquamarine4"
.
.
[646] "wheat" "wheat1" "wheat2"
[649] "wheat3" "wheat4" "whitesmoke"
[652] "yellow" "yellow1" "yellow2"
[655] "yellow3" "yellow4" "yellowgreen"
The way to select a particular color is:
plot(x,y,col=color()[646])
or
plot(x,y,col="wheat")
Ciao!
Subscribe to:
Posts (Atom)




