Wednesday, 11 March 2020

Coronavirus Data Analysis: Very Worrying Findings

I have been digging into the coronavirus data. First of all, the easiest one to access: Worldometer,
https://www.worldometers.info/coronavirus/
A useful data source, although it only gives historic data for the total case count and the death count. In their raw form, these accumulating totals are not very informative, so I turned them into a "counts per day" form and then plotted them on a log scale. Here is the cases per day graph.



The useful thing about using a log scale is that exponential increases (a sign of uncontrolled growth) show up as straight lines. On the left side of the graph you can see exponential growth that took place in China, early in the epidemic. Then, there is a dip, as China brought their outbreak under control. On the right. On the right, we see case rate increase again as the virus takes off in Europe. I was initially reassured that this growth seems to be slowing down. But then I took a look at the death rate graph.

The notable feature is that the death rate seems to be accelerating, if anything. It is now at its highest rate ever, surpassing the peak of the Chinese wave. Was this discrepency just a glitch?

To understand more, I downloaded the Johns Hopkins University dataset from GitHub. This seems to be the best source of simple, QC'd data on Covid-19. I scratched out some Python code to do similar plots, broken down by country, and focusing on a few places of interest; especially Hubei, North Korea, Iran and some key European countries. Below are the case rate and death rate plots:

Look at the death rate graph... Most noticeable is the exponential growth occurring in Italy. Also in Iraq, although at a slightly lower rate. This shows that the epidemic is completely out of control in these countries. Now look back at the new case count graph. The growth for these countries is tailing off. So I did a cross check of deaths reported against total cases. Below is the result.

The ratio of deaths to true cases (i.e. the fatality rate) SHOULD be relatively constant. So, a high apparent death rate indicates a very poor rate of detecting cases. So, what this shows is that Italy has exponential increase in cases and deaths, but a terrible, and worsening, case detection rate via testing., almost ten times worse than the best performers. Iran are at least slightly mire under control. The UK is doing better, and Korea, who seem to have things under control now, have done best of all.

The lesson: Beware Italy and Iran. They have immense problems and no solution in sight.


Thursday, 5 March 2020

Coronavirus - why politicians and journalists should learn multiplication

Context switch. Coronavirus. Covid-19. Something tells me I am going to be blogging about it a lot in the next few months. And that something is data science and models. Disease spread models tell us, based on some assumptions, what is likely to happen next. In my other life, I have been playing with disease spread models a lot recently, under a research project, and have come to understand their ways.

Currently I am worried about Covid-19, and I believe you should be too. This is a case where being scared can save you. It can save us all. I'll explain all future posts, but here are the basic facts.

When the virus first emerged in China, it was not noticed for a while. When doctors did begin to notice an unusual increase in a particular type of pneumonia, the information was suppressed, for political reasons. Nobody did anything to slow the spread, so it spread like wildfire. Finally the Chinese government came to their senses and resolved to shut down the exponential growth in cases. They were remarkably successful, because they took it seriously, to the maximum degree. Had they not done so, the ever multiplying number of cases would have continued until most of the population had been infected. Multiplying is the key word here. Cases would have doubled every few days - doubling and doubling.

So, the Chinese got a grip, They know how to shut things down. It was remarkable.

But cases dispersed out from the core, travelling on aeroplanes to all corners of the world. And now, the growth begins again. In Iran, where the problem was not acknowledged and again there has been exponential growth. In Italy, it seems to have been spotted rather late, leading to another case of explosive growth. The lessons from China are very clear; If you cut pairwise contact down by an order of magnitude, the virus spread CAN be shut down, but the government here in the UK (for here I be|) is playing it another way. They will start to think seriously about shutting it down, once it gets big and scary. For now, they will watch and wait. This is the most dangerous nonsense. The course of the virus is already set; There are already hundreds of people incubating the virus in the UK. The growth will be exponential, and it will be harder and harder to stop, the longer this goes on.

Key messages for today:

  1. Take this seriously; it's coming
  2. Take the maximum precautions you can; minimize pairwise contact, wash your hands, eat food cooked at home, etc.
  3. Stock up on food
By avoiding catching this at all costs, you protect yourself and others.

I will talk about the models and what the data is saying in future posts. Sadly, journalists and politicians don't understand models and data science. They don't understand, so YOU need to. 


Tuesday, 5 March 2019

Bitcoin woes

Just veering off the main topic for a minute to visit old pastures. A couple of news stories about bitcoin exchanges and disappearing funds...

https://www.ccn.com/new-zealand-bitcoin-exchange-cryptopia-postpones-launch

https://www.independent.co.uk/life-style/gadgets-and-tech/news/bitcoin-exchange-quadrigacx-password-cryptocurrency-scam-a8763676.html

This is by no means the first time bitcoin has disappeared of course. Remember MtGox and Mark Karples, or the Silk Road investigation, where a huge amount of coins went missing during the investigation.
Something about these latest cases is a little mystifying to me. I did some work on tracing coin a few years back. Bitcoin is an open ledger. It was always pretty easy to track coins at any point, to find out when coins were transferred out of a wallet and follow the trail downstream, with the right tools. And analytics have improved plenty too, with techniques like heuristic clustering used to cluster wallets together by ownership (working out who owns a given wallet is the harder part). There are companies that specialize in bitcoin analytics. So, what's going on? Why is it so hard to follow illicit coins and blacklist them? Is there some nation state actor (I wonder who?) at work stealing it and converting to hard currency? Perhaps ransomware isn't raising enough these days?

Saturday, 9 February 2019

Called it right!

Liverpool 3, Bournemouth 0! Well done boys!

Its a little bit unfair when forecasters take credit for "getting it right" when they make a probability-based forecast.  Not sure if all my odds were right, but the maximum likelihood option did come in!

Liverpool v Bournemouth Today: Analytics and Form Heatmaps

I have been running some analysis on today's games. Here are some numbers for one of the crucial ones - Liverpool v Bournemouth. I've developed a "heatmap" to show the form of both teams!

Here is a heatmap for Livepool. Attcking from increases left to right, while defensive form increases up the page. Believe me, this shows phenomenal form!


Compare this with Bournemouth. They are clearly less good!!

So, here is my odds estimate!

    x:0  x:1  x:2  x:3  x:4  x:5  x:6  x:7  x:8  x:9
0:y  57/1 90/1 367/1 1723/1 8332/1 49999/1 --- --- --- --- 
1:y  15/1 28/1 90/1 426/1 2272/1 16666/1 49999/1 --- --- --- 
2:y  9/1 16/1 58/1 234/1 1999/1 49999/1 --- --- --- --- 
3:y  8/1 13/1 45/1 229/1 1922/1 16666/1 49999/1 --- --- --- 
4:y  9/1 15/1 58/1 242/1 2499/1 16666/1 --- --- --- --- 
5:y  13/1 22/1 78/1 367/1 2173/1 16666/1 --- --- --- --- 
6:y  22/1 39/1 123/1 609/1 4999/1 24999/1 --- --- --- --- 
7:y  47/1 75/1 235/1 1281/1 8332/1 --- --- --- --- --- 
8:y  105/1 179/1 514/1 7142/1 24999/1 --- --- --- --- --- 
9:y  164/1 275/1 745/1 4999/1 49999/1 49999/1 --- --- --- --- 

summary:  Home win, 1/8; Away win, 29/1; Score draw 17/1; No score draw, 57/1.

Looks like the most likely result (by a whisker) is 3 - nil to Liverpool.


A reminder, this is a new method and still needs to be validated, so use with caution.

Wednesday, 6 February 2019

Everton v Man City Tonight!

So, here is my first odds forecast, for tonight's prem game...

Here is a grid of the odds for all score combinations up to 9 apiece...

       x:0     x:1  x:2   x:3  x:4  x:5  x:6  x:7  x:8  x:9 
0:y  19/1   8/1  10/1 13/1 32/1 74/1 207/1 768/1 2499/1 9999/1  
1:y  17/1   9/1  9/1 15/1 28/1 70/1 226/1 832/1 1999/1 9999/1  
2:y  42/1   19/1 19/1 31/1 68/1 146/1 285/1 1666/1 - 9999/1  
3:y  151/1 59/1 61/1 92/1 166/1 369/1 1110/1 - 9999/1 -  
4:y  344/1 262/1 178/1 434/1 999/1 1999/1 3332/1 4999/1 - -  
5:y  2499/1 1110/1 908/1 1666/1 3332/1 4999/1 - - - -  
6:y  4999/1 4999/1 - 9999/1 - - - - - -  
7:y  9999/1 - - - - - - - - -  
8:y  - - - - - - - - - -  
9:y  - - - - - - - - - -  


And here are some summary odds: 
Home win: '44/10', 
Away win: '1/2', 
Score draw: '5/1', 
No score draw '19/1'

These are based on 10,000 modeled games and 5000 particles per team. Interested to know if there are other odds people would be interested in.

I compared with Bet365 odds, which are generally fairly similar. My numbers seem to like the idea of a home win slightly more than Bet365. 1 nil to Everton looks like a value bet (although odds are long).

Health warning: I am still validating this model, although I believe the approach is generally solid!

*** Update*** the game finished 0-2 (win for Man City). It stood at 0-1 until injury time, which would have tallied with my most likely result

Sunday, 3 February 2019

Back to Bayes-ics

As explained last post, to do our football analytics, what we need are some input parameters about how "good" the two teams facing each other in a match are likely to be, on the day. There are two alternative approaches to doing this. One is based on classical statistics. To follow this approach you look back over a load of matches and work out an average scoring rate and an average rate of conceding. You can also estimate, on average, how much better the team performs at home. This approach has some weaknesses though. A team can get better or worse over the season; Its no good at telling how good the team will be today. It also requires quite a lot of data (a lot of matches) and assumes all the teams form are stable over time. In other words, it makes some assumptions that are not true. Which is never good.
A much better approach is to use Bayesian statistics. Thomas Bayes was a statistician with a keen interest in games of chance. Hence his work is very relevant to all sorts of gambling! The formulas he gave us are all about inferring the underlying truth from a series of observations. Each observation modifies our belief in a given hypothesis. To cut a long story short, bayesian inference crops up everywhere, in modern analytics.
The particular method I am deploying for football match analysis is the Particle Filter - a modern development, based entirely on Bayesian inference. You can find a pretty good intro to particle filters in this slide deck. Note the reference to football results analysis on slide 24... Using a PF for football analysis is s a nice party trick that often crops up in tutorial material, although I do it in a slightly more sophisticated way to the standard approach.
Applying a particle filter to the English Premier League works like this:

Each team is represented by a large number of "particles", each of which is a guess at the "model" - i.e. the qualities of the team (its attacking strength, defensive strength etc.)
  • Between fixtures, we "advance" these models, saying in effect. Last week the team was like this, so this week, how might the team have moved on
  • After a fixture, we "filter" the particles, preferentially keeping those that best explain the result. Incidentally, this is where Bayes comes in. His theorem says that instead of asking the hard question, "how good is my particle (model), given the result", we can ask "how likely is my result, given my model". This turns out to be an easier question and one we can answer. Importantly we consider not just the result but the capabilities of the teams involved. Hence all the analysis is interconnected.
  • Now, when two teams face off, we have a set of guesses about the teams capabilities at the present time that is based on all previous results, especially the last result. We can model the game considering the full range of guesses and get the best possible odds prediction, given the evidence.
In a nutshell, that's it. My plan now is to publish some predictions before the weekend fixtures and try to ascertain if we can beat the bookies!. That's my goal. Bookies are there to be beaten after all.

Saturday, 2 February 2019

Monte Carlo Football Analytics - Project Monaco

I'm going to christen this effort Project Monaco. It's about Monte Carlo and Football, so it's a no brainer, right?

I'll explain the basics of the analysis...

Consider two football teams, facing each other. The result is based on a few things. How good is the home team at attacking? Conversely, how good is the away team at defending? These two things will help determine an average rate of goal scoring for the home team. There is another factor in there - the home team advantage. Some teams perform better at home than away. Others do not, and some teams even do a little better away from home (home fans can be off-putting if they're not getting behind the team). On the flip side, how good is the away side at attacking, and how well can the home team defend.
In my method, I put these numbers into a pot, and work out an average expected rate of goal scoring for each team, in the context of them playing each other. I then create a computer model of the fixture, and run about ten thousand trial games, recording the result of each. What I get is a comprehensive odds forecast, covering every score permutation.
The computer model uses a classic statistical method to model the results of each trial game. The binomial distribution. Its hard to argue with the basics of this. The one hard part we are left with is establishing the input data for the match, regarding the strengths and weakenesses of each team. The truth is, we can't be certain about them, and that leads to some complications. What we need to do is consider a range of possibilities regarding this input data, which makes things a little mroe complicated. I will explain how we work out these inputs based on the league results in the next post!

Context Switch - Monte Carlo Football Analytics

It's been a year since m,y last post on Monte Carlo modelling of horse races. The reason for the gap is that I realised, through quite a bit of experimentation, that my methods for predicting the "true odds" of a horse race really weren't fit for purpose. A lot of the time, my odds predictions were quite similar to those from the bookies. But the cases where I predicted a horse should have much shorter odds (i.e. the "value bets" did not win as often as I expected. Essentially there was very little profit to be had. None, in fact, beyond random outbreaks of good luck.
I concluded that, for horse racing, there was a lot of information out there that helps understand how well a horse is likely to perform. Maybe a lot of it is on "back channels" known only to the inner racing community.  But, I concluded, at a minimum, you really needed to look back at a horse's history and see how it fared against each horse it raced, and to know how good each of these other horses were. That, I concluded, was too much like hard work (at least for now). Too much bespoke web scraping to be written for one thing, and life is too short. What I did do, is dream up an analysis method that could genuinely work. Its based on established methods, although it has elements that I don't believe have been tried before. But I decided it would be much easier to operate this on a more restricted field of runners. Like a football league, where a small set of clubs face each other in a very well defined, exhaustive set of fixtures. The Premier League, will be my case study!
The good news is, having tried this, I KNOW I have a method that is at least pretty cool. As I planned to do with the horse racing, I will provide more details about the method, and some of my odds predictions for upcoming games. More posts to follow!

Tuesday, 20 February 2018

Monte Carlo or bust?


Looking around a bit shows I am not unique in using Monte Carlo simulation to estimate the  true odds, in horse racing. For examole, see this monte carlo example. The author has a nice idea of using standard horse ratings, namely Official Rating (OR) and the Racing Post Rating (RPR) as the input to the model. My method is to use the past performance of each horse as the input - which I believe has some neat benefits - but the basic approach is the same.

Basically the first we have to do is define a probability density function (PDF) for the speed we think each horse might run in the race. It might look like this:

It represents the probability the horse will run at any given speed. The peak of the PDF represents the most likely speed for the horse. It may run faster or slower, but each are less likely. The PDF tails off at the edges to show this. The extreme edges are getting pretty unlikely. The shape of the PDF is important. The typical thing to do is to use a Normal or in other words Gaussian form for the PDF. This is not a bad choice, because Gaussian PDFs crop up all over the place in nature, so the likely running speed for a horse probably follows one.

When we execute the Monte Carlo race model we run a lot of imaginary races (maybe 1000 or more) and simply count the times each horse wins. To simulate each race, we draw, at random, example speeds for each horse from the PDF- such that the most likely race speed for each horse is in the middle of its distribution, with the frequency falling off towards the edge of the distribution We then rank the horses based on speed and work out the winner. Here the choice of the Gaussian PDF is handy because computer languages often have a ready made function for generating random samples from a nor distribution. I'm using Python, and it does the job nicely.

So, we have the results of a thousand or so simulated races, so now we can calculate the "true" odds for the horses easily enough. We made a few assumptions along the way, but if these are true, our odds should be good. Just to state those assumptions again:

  • We assume the past performance of the horse provides a good measure of its quality
  • We assume the horse's speed PDF is a Gaussian distribution positioned in proportion to the horse's quality score. 
  • We have to "invent" a width for this Gaussian, (called the standard deviation). This is a bit of a weakness, in that we have to make this up to make the odds look right. Still, it should probably be fairly constant for all races.
So, that's the basis of the approach. I'm working on some improvements that are quite subtle, and I'll introduce in future posts. Now lets see how well it works!

Thursday, 15 February 2018

Horse Race Analytics Explained

As discussed in the previous post, the aim of my horse race analytics is to estimate the "true odds" for each horse in a race, so that we can compare them with what bookmakers are offering, to identify "good value". I put the term "true odds" in quotes (I did it again!) for a reason. It is quite hard to say what the true odds of anything is - let alone a one-off thing like a race. To measure the true odds accurately we would need to run the same race, with the same horses,  the same health, in the same weather and track conditions, a very large number of times and count the outcomes. But the weather is actually an uncertain factor so maybe we'd need to use a few different seasonally appropriate weather conditions. Anyway, clearly this is completely unfeasible, and the true odds are therefore really quite an abstract concept. Its doubtful that a precise number value for the true odds can even be defined in fact. One thing we can do is work out our way of calculating true odds and test it over a large number of real races. Our three to one (3/1) horses should on average win once for every three losses, our 2/1 should win once for every two losses etc. Incidentally if you are not familiar already, it turns out odds are a little different to probabilities. They are quoted as the number of losses vs. the number of wins. The win probability on the other hand will be a number between 0 and 1 defining the likelihood of a win in each race.

My approach to calculating the true odds has two stages. First, we estimate a "quality score" for each horse. The better the horse, the higher the quality score. We need to decide a standard way of doing this - for example we could look at the six previous races and assign a score of 5 for a win, 4 for second etc. I have a more sophisticated approach that I'll describe in a future post, but for now, you get the general ideas. 

Having calculated a quality score for each horse, we then use this to estimate the odds. In simple, made up cases, this might be easy. If three horses race, for example, and they are all exactly equal in terms of quality score then the true odds are 2/1 for each horse, because, on average, in each 3 races, each horse will win one and loose twice. For any significantly complicated example, it gets a lot more difficult. In fact, it rapidly becomes quite impossible to calculate analytically (i.e. using a formula). The way to do it is to do the kind of thing  merchant banks do a lot when analysing trades - Monte Carlo modelling. More about how this works in the next post.

Monday, 12 February 2018

Idle Pursuits

I decided to talk about my latest hobby - horse race analytics. Its something I started thinking about twenty years ago and, after several false starts, I believe I have finally worked out the maths of what I wanted to do and coded up an approach that works. I'm not sure why it took so long!

The basics. Horse racing is an uncertain business. Generally speaking it is never possible to accurately predict the results of a horse race, unless you are a) veeeery lucky or b) own all the horses. The best prediction is usually that the favorite will win. But typically the odds the bookmakers will give you on that happening will not be very good, so if you do it every time, you will, in the long run, loose money. If you back the horse where the bookies are giving the best odds, i.e. they pay you the most for a win, you will also loose money overall, because these horses will win less often. Not never, just less often.

There is, surprisingly, one reliable strategy for making money form horse racing, and that is to pick "value winners" which means that you pick horses where the bookmakers are offering "good value". In other words, they are offering better odds than the quality of the horse would suggest. In yet more other words the horse is more likely to win than they think it is. The "true odds" of the horse are "shorter" than the bookies odds. So, all we have to do is work out the true odds and back horses where the bookies are offering longer odds.

Therein, of course, lies the problem. How to calculate the "true odds". How to calculate odds better than the bookies, whose job it is to do this. They have teams devoted to it; observing races, going to stables, timing, observing, etc. This is what I'm trying to do. It won't be easy - but I think I at least have some maths that can help. I'm planning to use the kind of stuff economists and stock traders know about (at least some of them). To drop in a name, Bayes is the key to this. Bayes is the key to a lot of things.

What I plan to do is develop and refine the method over the coming weeks, publishing what I calculate as the true odds, comparing these to the bookmakers odds and highlighting my betting  recommendation. Over the weeks, we'll see if its working and hopefully refine as we go!

Thursday, 8 February 2018

Bitcoin analytics - reprise

I saw this twitter post from Justin Seitz @jms_dot_py the other day:

https://twitter.com/jms_dot_py/status/958741474572750848

Its good to see this push to uncover the underside of the crypto currency ecosystem going forward.  It again mentions the point that the blockchain is a public ledger, which lists all accounts and all transactions that have ever occurred - forever. That gives quite a bit of time for us to pick over it. De-anonomizing accounts is the only barrier to getting full context on those transactions. Of course, there is only so far it can be taken right now, but in future, with better tools, who knows what we will uncover.

I first came across Justin Seitz through his Black Hat Python book - which is packed full of ingenuity. He also does a lot in the OSINT space and developed the very good Hunchley tool for OSINT investigations. Definitely not a black-hat, but he does like to look under stones. I follow his Hunchley daily dark web report - which is an example of that.

Thursday, 9 October 2014

MtGox - Time Line

A couple of days ago I dusted off some bitcoin trace data I produced a few months back and decided to look at other ways of processing it. I'm interested in what time correlation could tell us about the bitcoin transaction ecosystem. As a very quick starting point I decided to time-bin the transactions following the MtGox bitcoin show (see my previous post).

I modified my python post-processing routine to count the coin flow and the number of transactions in each of my chosen large time steps (5000 seconds) working from the very first transaction after the show point. Doing this I felt a bit like an astronomer analysing the ancient light from just after the big bang (ridiculous thought). Its pretty cool that this data is frozen forever for us to examine.

Here's some graphs I made by exporting csv and plotting in excel...


Here's the first graph - number of transactions per step. You can see the rate ramp up from near zero in the first step up to hunderds of transactions per step. Note, the x axis is time step number. The steps are 5000 seconds - around an hour and a half (daft choice - I should have used hours shouldn't I :?)



If we look at the value transfer rate (above graph) we see a different story. As time goes on, the actual values transferred fall sharply over time. That's because the transations themselves are getting smaller as the coins get more and more split up. Note the log scale!


So by dividing the volume by the transaction rate we get the mean size - this falls form the initial HUGE transaction steeply down. Note, I'm not considering dilution as I'm not quite sure how to handle it. probably  mean size would continue to fall over time if I looked at the amount of coin that really originated from the original pot as some of the later transactions might be getting quite dilute.

Here is the interesting bit though. I looked again at the first graph and saw what seems to be an outlier... The big spike near the start. So I replotted the very early parts - here's what I got...



There's a huge spike of transactions on around timestep 28. A leap up to nearly 900 transactions and then it drops back again. Clearly something happened here. What I'm thinking is that someone ran a script to disperse the coins widely. This probably makes the trail really hard to follow if you do it manually. Fortunately we have computers so its easy!

Now what I'm wondering if this spike corresponds to the strange whirlpool structure on my Gephi coin trace... I think that whole structure may have been formed very quickly by running a script. I'm now trying to imagine what the script code to do that would have looked like and exactly why it was done in the way it was. Next step is to add a time line to the Gephi trace and step through it. That'll let me see when that structure appears on the graph view. I will keep you posted :) 

Friday, 25 July 2014

The Open Source Personal Meme Filter

Memes to avoid:
Memes that people are out to get us
Memes that sections of society are out to get us
Memes that we (the guys in the white hats) are superior to others
Memes that the people of other countries are inferior, dislike us, are plotting against us
Memes that being unkind to people is fun or cool
Memes that its them or us
Memes that to attack the meme is wrong, that to question a meme is wrong, or reject a meme is wrong
Memes that society has been bad to us and we have a right to take revenge
Memes that deny us our free will - our right to choose our path
Memes that justify the unjustifiable

Fear memes, superiority memes, memes that surreder our free will. memes with their own meme defences. Complexes of mutually supporting memes with no substance (houses of meme cards)

In short - memes that bring us fear and give us licence to do bad things to people

Rules: accept only memes that are free standing, testable, cause nobody any problems, bring good to people, and most of all that you believe in

A further observation: This personal meme filter is itself a meme - hopefully one that passes its own acceptance rules. Its okay to attack this meme - but hopefully it is strong.

An open source meme: This meme is published under an LGPL licence. Please feel free to build on it, but do share your results with us all. Together maybe we can build a better world.

Wednesday, 16 July 2014

Bitcoin - Digital Fortress?

I just thought I'd float this idea: In his science fiction novel of a few years ago Digital Fortress, Dan Brown describes a supercomputer built by the NSA to crack complex codes like public-private key encryption. We know now from the Snowden revelations that they've been doing a lot of funky stuff - some involving code cracking - some more focussed on back dooring, intercepting and hacking.

But Bitcoin - this technology of unknown origin... It involves an ever increasing amount of computer power computing hashes at an exponentially increasing rate. Vast bespoke computers of unprecedented power are now dedicated to this. Its driven phenomenal advances in bespoke computer power - pushing advances in GPU and bespoke ASIC hardware. I guess the two questions that spring to mind are:
1) Who is really behind bitcoin? Anarchists? Bankers? Con Men? Aliens?... The NSA???
2) Could we have helped (lured by a few riches) build a vast Digital Fortress style global computer to help our government security guys crack the most complex codes. I dunno - is this eveb possible?

Wednesday, 2 July 2014

The Meme Meme

Thought for the day: The term meme was coined by Richard Dawkins in his book The Selfish Gene. A meme is an idea. Ideas exist within the population (who are their hosts). The point he was making was that strong ideas live and multiply and spread - just like strong (or perhaps more aptly, fit for the current environment) genes.
The idea of a meme is, clearly, in itself a meme. the meme meme (or meme2).
Interestingly common use of the term deviates from the original. It has evolved into something new. The meaning in most quarters is now this. Grotesque image plus big white writing plus bad grammar equals meme. Simples.
But the real meaning is far more interesting. Currently our world is populated by some powerful and dangerous memes. I won't put specific names to them here. There are several - and they generally follow the same pattern. "I believe that my set of ideas is the only correct one - believers in all other sets of ideas are therefore by definition wrong". I disagree with this wholeheartedly. Many sets of ideas can get you through life very nicely and do nobody any harm, and hopefully do some good. There are only a few no-no's in my view, but complete intolerance of a diversity of views, cultures, physical characteristics etc. is definitely high on the list of bads. But why so many bad memes? Compare back against the gene example on which the meme meme is based and you immediately see the answer... Because they thrive in the current environment. That's the thing we've got to understand and do something about... What is it about the current environment that allows these quire damaging memes thrive? The environment is pretty much us, the human hosts, or human society. We need to take a serious look at that.

Tuesday, 1 July 2014

MtGox and Instawallet AGAIN!?

I really should move on from speculating about MtGox and the murky world surrounding it. But its HARD, with so much still obscure.

In a previous post I showed how Instawallet seems to show up prominently on the transaction trail after the MtGox dog-and-pony-show (by which I mean their "show of coins" back when). Anyway, I just found this thread dating back to when Instawallet nearly folded and got taken over. An interesting piece of history on several fronts. But check out post #19 from Mystery Miner. He seems to be reporting some wierdnes whereby when he tired to use Instawallet to buy ganja on Silk Road, the ganja was actually purchased with apparently stolen coins, while his actual coins got used to purchase a "killer wet job". I have know idea what that is, but that aside it does seem he is reporting some kind of Gox-Instawallet fund confusion or interchangeabilit . although how reliable his evidence is though may be called into question by the fact that he may be judged from his post to be:

a) a self-styled shadowy enigma and international cyber-man-of-mystery
b) a self-proclaimed buyer of illegal drugs and frequenter of the deep web
c) maybe just a bit of a sleaze

... he may have actually been onto something there though!

Monday, 30 June 2014

BitIodine Review

I just found this online tool called BitIodine. I've just had a look - it claims to do clustering etc and I was quite excited they were doing something a bit like my Gephi stuff. But it seems to be simple text-based stuff. I'm not terribly impressed - surely someone must be doing something more exciting out there?

Wednesday, 4 June 2014

Bitcoin 2.0

Its amazing how fast things move in the new world that is crypto-currency (Year 0 was only 2009). A couple of months ago we had the collapse of MtGox and others, a few months earlier it was the Silk Road seizure. Now suddenly things are looking more mature (I guess a shake-out of the weakest links HAD to happen), the market is up (bitcoin was at $660 today) and some good innovations are going on.

One of my favourite examples of bitcoin doing good things for the world is BitPesa. This allows you to send money to anyone in Kenya (and other parts of east Africa) buy buying bitcoin and sending it through the BitPesa gateway. Basically you can send real money to anyone who uses the M-Pesa system - a neat pay-by-text and microfinance system run by Safaricom and Vodacom - the mobile phone giants from Kenya and Tanzania. Mobile phone usage is MASSIVE in East Africa and M-Pesa is really huge there (see article), so that's a lot of people you can send money to instantly. PitPesa is run by an American ex-pat living in Nairobi so its obvious they undertand the opportunity and what it could achieve. I hope more of this kind of thing happens as it could open up opportunities for a lot of people in a beautiful part of the world.

I found this link with a movie of the world's largest bitcoin mining corp. They're using big banks of ASICs, each controlled by a raspberry PI. The head guy says his profits are great, but I did the maths and worked out (using a standard bitcoin calculator) that if you buy one of the mining ASICs, at current retail price you'd be looking at a 3 year payback period. That's not much use unless you factor in some dramatic bitcoiin value rises (possible I guess). I realised on thinking further that bitcoin mining will always stabilise at a point where profit is zero for most people, as whenever a better cheaper technology come along, people will buy it until the point where the world's mining capacity pushes the mining difficulty up to the point where profits are zero. Hence, mere mortals sould not expect to make any money. To make money you need an edge - you have to buy the hardware first, and negotiate a big discount. You need to do a deal with the electricity supplier. In short you need to be BIG, like these guys... Or you have to cheat, and get a bot net.