| Andrew Cooke | Contents | Latest | RSS | Twitter | Previous | Next

C[omp]ute

Welcome to my blog, which was once a mailing list of the same name and is still generated by mail. Please reply via the "comment" links.

Always interested in offers/projects/new ideas. Eclectic experience in fields like: numerical computing; Python web; Java enterprise; functional languages; GPGPU; SQL databases; etc. Based in Santiago, Chile; telecommute worldwide. CV; email.

Personal Projects

Lepl parser for Python.

Colorless Green.

Photography around Santiago.

SVG experiment.

Professional Portfolio

Calibration of seismometers.

Data access via web services.

Cache rewrite.

Extending OpenSSH.

Last 100 entries

Sentidos Comunes (Chilean Online Magazine); Hilary Mantel: The Assassination of Margaret Thatcher - August 6th 1983; NSA Interceptng Gmail During Delivery; General IIR Filters; What's happening with Scala?; Interesting (But Largely Illegible) Typeface; Retiring Essentialism; Poorest in UK, Poorest in N Europe; I Want To Be A Redneck!; Reverse Racism; The Lost Art Of Nomography; IBM Data Center (Photo); Interesting Account Of Gamma Hack; The Most Interesting Audiophile In The World; How did the first world war actually end?; Ky - Restaurant Santiago; The Black Dork Lives!; The UN Requires Unaninmous Decisions; LPIR - Steganography in Practice; How I Am 6; Clear Explanation of Verizon / Level 3 / Netflix; Teenage Girls; Formalising NSA Attacks; Switching Brakes (Tektro Hydraulic); Naim NAP 100 (Power Amp); AKG 550 First Impressions; Facebook manipulates emotions (no really); Map Reduce "No Longer Used" At Google; Removing RAID metadata; New Bike (Good Bike Shop, Santiago Chile); Removing APE Tags in Linux; Compiling Python 3.0 With GCC 4.8; Maven is Amazing; Generating Docs from a GitHub Wiki; Modular Shelves; Bash Best Practices; Good Emergency Gasfiter (Santiago, Chile); Readings in Recent Architecture; Roger Casement; Integrated Information Theory (Or Not); Possibly undefined macro AC_ENABLE_SHARED; Update on Charges; Sunburst Visualisation; Spectral Embeddings (Distances -> Coordinates); Introduction to Causality; Filtering To Help Colour-Blindness; ASUS 1015E-DS02 Too; Ready Player One; Writing Clear, Fast Julia Code; List of LatAm Novels; Running (for women); Building a Jenkins Plugin and a Jar (for Command Line use); Headphone Test Recordings; Causal Consistency; The Quest for Randomness; Chat Wars; Real-life Financial Co Without ACID Database...; Flexible Muscle-Based Locomotion for Bipedal Creatures; SQL Performance Explained; The Little Manual of API Design; Multiple Word Sizes; CRC - Next Steps; FizzBuzz; Update on CRCs; Decent Links / Discussion Community; Automated Reasoning About LLVM Optimizations and Undefined Behavior; A Painless Guide To CRC Error Detection Algorithms; Tests in Julia; Dave Eggers: what's so funny about peace, love and Starship?; Cello - High Level C Programming; autoreconf needs tar; Will Self Goes To Heathrow; Top 5 BioInformatics Papers; Vasovagal Response; Good Food in Vina; Chilean Drug Criminals Use Subsitution Cipher; Adrenaline; Stiglitz on the Impact of Technology; Why Not; How I Am 5; Lenovo X240 OpenSuse 13.1; NSA and GCHQ - Psychological Trolls; Finite Fields in Julia (Defining Your Own Number Type); Julian Assange; Starting Qemu on OpenSuse; Noisy GAs/TMs; Venezuela; Reinstalling GRUB with EFI; Instructions For Disabling KDE Indexing; Evolving Speakers; Changing Salt Size in Simple Crypt 3.0.0; Logarithmic Map (Moved); More Info; Words Found in Voynich Manuscript; An Inventory Of 3D Space-Filling Curves; Foxes Using Magnetic Fields To Hunt; 5 Rounds RC5 No Rotation; JP Morgan and Madoff; Ori - Secure, Distributed File System; Physical Unclonable Functions (PUFs); Prejudice on Reddit

© 2006-2013 Andrew Cooke (site) / post authors (content).

Raid 5 Speeds

From: "andrew cooke" <andrew@...>

Date: Sat, 7 Apr 2007 09:59:57 -0400 (CLT)

I finally understand (I think) how having raid 5 affects the disk speed.

For large inputs and outputs, and for seeks, it is much faster than single
disks, increasing in speed by a factor of N (or perhaps N-1).  You can see
this by running iostat while the disks are in heavy use.  The IO rates for
the md device are much higher than the disks themselves, and the disks are
limiting the system (lots of IO wait in the CPUs).

This is also clear in bonnie++ output:

Version 1.01d       -Sequential Output- -Seq-Inp- --Random-
                    --Block-- -Rewrite- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP  /sec %CP
quiet            7G 34986  10 15030   3 43503   4  73.6   0

(reformatted to drop the per character values)

For my cheap disks (320Gb SATA 7200 WD Caviar), those are significant
improvements over a disk in isolation, as you'd expect from striping.

However, "everyone" says how bad raid 5 is for writing, so why are the
write speeds so close to read?  Because, I think, the file being written
is much bigger than the 128K block size, so there's no need to read the
parity information when writing - the system is simply writing a new
parity block along with new data blocks.

This may be reflected in the lower "rewrite" figure, which is dirtying
data within a single block.

Andrew

Rubbish

From: "andrew cooke" <andrew@...>

Date: Sat, 7 Apr 2007 17:35:22 -0400 (CLT)

I think everything I wrote above may be rubbish.  I get the same numbers
on my laptop.  The test used a 4GB file and my current best guess is that
2GB stayed in memory, hence the double speed for read/write over rewrite. 
But really I don't have a clue.  And seeks were actually slightly better
on my laptop.

(The laptop feels slower, but that's more a CPU than a disk thing, I thin
- one reason I have been trying to measure/understand things is because I
sometimes feel there's an odd lag with the server's disks).

Andrew

Iozone Results (Disk Speed Again)

From: "andrew cooke" <andrew@...>

Date: Thu, 12 Apr 2007 21:07:08 -0400 (CLT)

I ran iozone overnight and the results for random reads are shown here -
http://www.acooke.org/cgi/photo.py?start=disk&cols=5&rows=3

I'm not sure how clear that is (I shrank the image a little to fit my
standard "photo" size), but file size is increasing towards the viewer and
throughput drops off a cliff when the file size exceeds memory (the drop
from orange to blue).  The two largest file sizes were 4 and 8 GB.

Along the "foot" of the plot is record size.  There's a slight dip when
record size exceeds about 1MB.  This might be related to the L2 cache (2MB
shared).

iozone - http://www.iozone.org/

Also, gnuplot in X now lets you "drag" 3D plots with the cursor.  Very
cool + useful.

gnuplot - http://www.gnuplot.info/

Andrew

Comment on this post