Mittwoch, 6. April 2016

Integrate raspberry pi 3 into spark cluster

I have the spark cluster already set up, so this is only about setting up a fresh raspberry pi 3 as a worker in the cluster.

This is based on setting-up-a-standalone-apache-spark-cluster-of-raspberry-pi-2

Installing raspbian
Put the sd card (I use a 16 GB Toshiba ADP-HS02) in your card reader.
I then use lsblk to find out the mount path of the sd card:

In my case it's sdb1, so that (sdb) is where I want to put the raspbian image:

sudo dd bs=1M if=~/Downloads/2016-02-09-raspbian-jessie.img of=/dev/sdb

This can take some time. When it's done unmount the sd card.

Physically attaching the pi to the cluster

Put the sd card in your pi.
Connect the lan port to the switch.
Power up the pi.

Configuring the pi
ssh into the headnode/master of your cluster (assuming it is a standalone cluster). I gave it the hostname "pi-headnode", so I can access it via:
ssh pi@pi-headnode
You can use nmap to scan the cluster for nodes.
As a reminder: This is a standalone cluster. So it has its own network, with its own IP addresses. The headnode is the bridge between the cluster, and the main network. From the headnode, both networks (the main network and the cluster) can be accessed. Use ifconfig to find out the IP addresses of the networks. eth0 will probably be connected to the cluster, and eth1 to the main network, but this depends on your setup. In my case, the IP of the main network is 192.168.0.*, that of the cluster 192.168.1.*. So to scan the cluster I use:
nmap -sn 192.168.1.0/24
and I see the new node.

run the scripts
sh map-network
sh copy-keys
sh upgrade-workers
to integrate the pi into the cluster and upgrade all pis. This can take some time.
Instead of upgrade-workers, you could ssh into the new pi and upgrade it manually with sudo apt-get upgrade.

The pi is now part of the cluster, but I like to configure it a bit more:
ssh into the new pi
start the configuration tool:
sudo raspi-config
and expand the disc space, set boot to console and automatic login to save ram, reduce ram for GPU to 0

copy spark to the cluster
update the slaves.txt
cp ~/workers.txt ~/spark-1.6.1-bin-hadoop2.6/conf/slaves
copy spark to the cluster
scp -r /home/pi/spark-1.6.1-bin-hadoop2.6/ pi@192.168.1.12:/home/pi/spark-1.6.1-bin-hadoop2.6/

 
Start the slaves
spark-1.6.1-bin-hadoop2.6/sbin/start-slaves.sh
this starts all slaves that are not yet running.

Done!

You will hopefully see something like this on your webUI





Mittwoch, 16. März 2016

free disc space in /boot by removing old kernels

Once in a while Ubuntu can't update, because there is not enough free space in /boot.
 
To free space, old, unused kernels have to be removed. This can automatically be done with: 
 
 
sudo apt-get purge $(dpkg -l linux-{image,headers}-"[0-9]*" | awk '/ii/{print $2}' | grep -ve "$(uname -r | sed -r 's/-[a-z]+//')")
 
From:
 
https://askubuntu.com/questions/89710/how-do-i-free-up-more-space-in-boot 
 

Sonntag, 21. Februar 2016

configuring raspbian jessy for easy ssh access


Configure your pi with
sudo raspi-config
Within the menu, do the following
- change the name of your pi, for example 'pi1'
- allow ssh access
- change the password (and remember it)

restart the pi

Now, from some other unix computer, you want to scan the local network for connected devices. This can be done with nmap. You also need to know the IP address of your local network, which is normally 192.168.0 or 192.168.1, but better find it out with either ifconfig or ip addr.
Step by step:
- Install nmap
sudo apt-get install nmap
- find your local networks IP address with either
ip addr
or
ifconfig

joker@joker-Ultrabook:~$ ifconfig
eth0      Link encap:Ethernet  HWaddr e8:03:9a:ee:f4:2b 
          inet addr:192.168.0.131  Bcast:192.168.0.255  Mask:255.255.255.0
 
In this example, my devices IP is 192.168.0.131, so my local networks address is 192.168.0, 131 being my laptops number.

- search the network with nmap:
nmap 192.168.0.0/24

24 is shorthand for subnet mask 255.255.255.0

 As a result, nmap shows you all devices in the network, including your pi. Of course the pi has to be connect to the network via cable or wifi:
Nmap scan report for pi1 (192.168.0.119)
Host is up (0.011s latency).
Not shown: 999 closed ports
PORT   STATE SERVICE
22/tcp open  ssh

Nice, the pi is there, ssh is open, and its correctly named 'pi1'.
You could use its IP address 192.168.0.119 to ssh into it, but using the name is much easier.
ssh pi@pi1




 



Freitag, 19. Februar 2016

I checkout back into my master, but it still looks like the branch I was working on?

Short story:

You have to commit your changes in the branch before checking out to the master. So the solution is to checkout back to the branch, commit the changes, and checkout to master afterwards.


Long story:

I started implementing multiprocessing. Naturally, I didn't want to mess up my code, so I created a new branch 'multiprocessing':
git checkout -b multiprocessing
Some hours of biting my nails later, I had it running, and even saw some improvement in speed. I also learned, that implementing multiprocessing in python is easy, but finding the right tasks that can be processed in parallel with speed improvement is hard.
Anyway, the code I wrote was far from clean, but my curiosity was satisfied for now. I  decided to put that project aside and get back to it later.
So I checked out to my master:
git checkout master
When I ran the tool later, I realized that it started working on multiple cores. Had I forgotten to checkout? So I had a look:
git branch
and it confirmed that I had correctly checked out:
joker@joker-Ultrabook:~/VCF2AFAnalysis$ git branch
* master
  multiprocessing
  plot_labels

Still, the code was the one I wrote while checked out to multiprocessing.
So I checked out back to multiprocessing:
git checkout multiprocessing
and commited the changes
git commit -am 'started implementing mp'
After checking out to master everything was alright.


I don't know why git doesn't tell you that you always have to commit your changes before checkout, but well, that's just how he wants it.

Freitag, 22. Januar 2016

How to qrsh into a specific host

Working on a grid engine you sometimes want to work on that one specific host, using qrsh instead of ssh.

Somehow it is not super obvious how to get there, and even google didn't help.

So there:

qrsh -l hostname=waldorf
where "waldorf" is the name of the host you want to log onto.

Freitag, 15. Januar 2016

How to prevent "broken pipe" with ssh (Ubuntu 14.10)

Since I work mostly from home, I use ssh a lot.
The "broken pipe" disrupted my workflow continuously, esp.
when some task running for a couple of minutes was killed.

At first, I fixed it in a rather n00bish way, by running the task in the background
('&' or (ctrl+z and bg)), and keeping ssh alive with a never ending while loop:
while true; do sleep 60; qstat -u ries;  done
But according to
https://askubuntu.com/questions/127369/how-to-prevent-write-failed-broken-pipe-on-ssh-connection
there are of course better ways:

send a keepalive message to the server

This is basically the same idea as with the while loop. Every 60 sec or so, a message is send to the server, telling it to keep your ssh connection stable.

edit your
~/.ssh/ssh_config 
with:

Host *
ServerAliveInterval 60 



use the screen shell for long-running tasks


For tasks running longer than a couple of minutes (in my case it can be days), it doesn't really make much sense to keep a stable ssh connection, and your hoe computer up and running.

A better solution is to start the task in a way, that you can detach it for now, and get back to it later, while it goes on on it's own. Probably the best way to deal with this is screen, although it takes some time to get used to it.


Screen is a powerful utility that allows you to control multiple terminals which will stay alive independently of the ssh session.

While logged in via ssh, start screen


screen

you are now in a new screen shell. Start your command. For me s.th. like this:



python ../bwa_parallel_arrays/bwa_parallel.py -P tmp -I . -t 4 -e error -o out -n 10 > screen_out.txt &

detach the screen shell and get back to your previous terminal:


screen -d



Detach the screen session (disconnect it from the terminal and put it into the background). A detached screen can be resumed by invoking screen with the -r option.

screen and your task keep running on the host you logged on, as can be seen with the top command:




PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
55557 ries 20 0 3948712 3.686g 6392 R 5.9 1.5 0:41.99 python
32063 ries 20 0 27036 3524 2284 R 1.3 0.0 0:00.23 top
9023 ries 20 0 28976 3076 2440 S 0.0 0.0 0:00.00 screen


reattach your screen session with


screen -r

You can also check the status of your (various) screen sessions with
screen -list

There is a screen on:
9023.pts-8.waldorf (Detached)
1 Socket in /var/run/screen/S-ries.


using 'disown'


Out of convenience, I tend to just start the tasks in the background with '&', and detaching it from the shell with 'disown', so it doesn't get killed when the shell gets killed.

This is very easy, and can be done even after starting the task, by sending it to sleep background with Ctrl + Z, then running in the background with 'bg' and detaching it with 'disown'. This somehow doesn't work well with ssh. When the pipe breaks, the task breaks.


using 'nohup' [update]

Better than using 'disown' is to use 'nohup' (no hangup).
Just prepend it to your command, and it will not terminate if the terminal gets closed.




Dienstag, 5. Januar 2016

GitHub: pushing local changes to the remote version (updated)

After fumbling around for a while, I finally got a grip on how to use GitHub.
If you're as slow a learner as me, it might take you a couple of days to really get used to it, but it's also a skill worthwhile to learn.

For now, I'm just using github for remote access, and version control for my scripts.

I assume, you already have a github project initiated.

The best in-depth explanation I could find:
https://www.atlassian.com/git/tutorials/comparing-workflows/centralized-workflow

 

Task:

Editing some file of the program locally, then merge it with the remote version.


Get a local copy of your own rep
git clone https://github.com/davidries84/bwa_parallel_arrays.git

Edit some files.

If it is your own rep in which you changed files and you just want to update the master:

git commit -am "change time tracking"
-a means all files you changed will be pushed to the remote repository
-m is used to state the changes you did in a short message

By 'pushing' the changes up to the remote version of your code, you update that remote version:
git push -u origin master




If you added some files, instead of only editing already existing files:
git add .
git commit -am "adding test files"git push -u origin master



Editing the 'master' should only be done for failsave changes (like comments). The master should always be fully functional code. So for implementing new features or bugfixes, you should use 'branches', which get merged into the master, when the new code is ready and finished.


In your local directory, create a new branch for your intended changes.
By creating a new branch, you automatically change (checkout) into that branch.

git checkout -b update_Readme


Do your changes, i.e.
gedit README.md

"commit" your changes:

 from the man "git commit" pages:
Stores the current contents of the index in a new commit along with a log message from the user describing the changes.

The content to be added can be specified in several ways:
...
4. by using the -a switch with the commit command to automatically
"add" changes from all known files (i.e. all files that are already
listed in the index) and to automatically "rm" files in the index
that have been removed from the working tree, and then perform the
actual commit;

So we commit everything (-a) and add a message describing the change (-m)
git commit -am "added some text to readme"
To update the repository, we have to "push" the local branch "update_Readme", to the origin.
git push " Updates remote refs using local refs, while sending objects necessary to complete the given refs."

 git push origin update_Readme 


Afterwards, we can checkout back to the master.
git checkout master 

After accepting your own changes (using the website), it makes sense to sync your local master with your remote master, thus bringing it up to date.
git-fetch - Download objects and refs from another repository

git fetch upstream
git-rebase - Forward-port local commits to the updated upstream head
git rebase upstream/master