2020-06-28

Combinations and permutations

Counting is simple except when there is a lot to be counted. Combinations and permutations are such a case; they are about counting without replacement. Suppose we want to count the number of possible results we can obtain from picking k numbers, without replacement, from an equal or larger set of numbers, that is, from n where k \leq n. When the same set of numbers in different orders should be counted separately, then the count is called the number of permutations. So, if we have some set of numbers and shuffle some numbers around, then we say that the numbers are permuted. When the same set of numbers in different orders should be counted only once, then the count is called the number of combinations. Which makes sense since it is only about the combination of numbers and not the order.

Show more

2020-06-27

Comparing means and SDs

When comparing different papers it might be that the papers have numbers about the same thing, but that the numbers are on different scales. Forr example, many different questionnaires exists measuring the same constructs such as the NEO-PI and the BFI both measure the Big Five personality traits. Say, we want to compare reported means and standard deviations (SDs) for these questionnaires, which both use a Likert scale.

In this post, the equations to rescale reported means and standard deviations (SDs) to another scale are derived. Before that, an example is worked trough to get an intuition of the problem.

Show more

2020-05-11

Predicates and reproducibility

While reading texts on statistics and meta-science I kept noticing vagueness. For example, there seems to be half a dozen definitions of replicability in papers since 2016. In this text, I try to formalize the underlying structure.

Edit 2020-11-01: The model below is basically the same, but poorer, than the causal models as presented by, for example, Pearl (2009).

Assume determinism. Assume that for any function f there is a set of predicates, or context, C which need to hold for the function to hold, that is, return the correct answer. Let this be denoted by C \xRightarrow{a} f. For example, Bernoulli's equation solved for \rho only holds for a context C_b containing isentropic flows, that is, C_b \Rightarrow \text{Bernoulli's equation}, where C_b contains isentropic flows. There have been arguments that such contexts need to contain an (open-ended) list of negative conditions (Hoefer, 2003). Let these contexts and the contexts below also contain this list.

Show more

2020-03-05

Simple and binary regression

One of the most famous scientific discoveries was Newton's laws of motion. The laws allowed people to make predictions. For example, the acceleration for an object can be predicted given the applied force and the mass of the object. Making predictions remains a popular endeavor. This post explains the simplest way to predict an outcome for a new value, given a set of points.

To explain the concepts, data on apples and pears is generated. Underlying relations for the generated data are known. The known relations can be compared to the results from the regression.

Show more

2020-02-02

The greatest sales deck someone else has ever seen

According to Andy Raskin, the greatest sales deck has five elements. In this post, I'll present an adapted version. In line with the rest of this blog, I'll give an example of selling a programming language which is not Blub to a company, the newer language is called Y. Assume that the company is fine with language Blub because, well, everything is written in Blub and all the employees know Blub.

Show more

2020-01-24

Correlations

Correlations are ubiquitous. For example, news articles reporting that a research paper found no correlation between X and Y. Also, it is related to (in)dependence, which plays an important role in linear regression. This post will explain the Pearson correlation coefficient. The explanation is mainly based on the book by Hogg et al. (2018).

In the context of a book on mathematical statistics, certain variable names make sense. However, in this post, some variable names are changed to make the information more coherent. One convention which is adhered to is that single values are lowercase, and multiple values are capitalized. Furthermore, since in most empirical research we only need discrete statistics, the continuous versions of formulas are omitted.

Show more

2020-01-16

Benefits of writing blog posts

The first step into creating good habits is figuring out why exactly you want the habit. To me, writing blog posts seems like a good habit, but I'm unsure why. This post will attempt to convince the reader and myself of the benefits. I have combined my own ideas with the ideas by Terry Tao [1] and Gregory Gunderson [2], and grouped them.

Pedagogic benefits

  • Writing detailed expository notes is a way to practise research. This allows you to break free from the methods you are used to [2].

  • One can practise writing [1].

  • Writing allows one to test understanding of an idea [1]. It forces you to explain it clearly without hand waving. When aiming your text at colleagues or future employers, you cannot use jargon to hide your lack of knowledge [2].

  • Writing allows figuring out what exactly you do not understand or what you need to learn first [2].

  • Writing aids in structuring knowledge [2].

Show more

2019-12-29

Statistical power from scratch

In the 1970s the American government wanted to save fuel by allowing drivers to turn right at a red light (Reinhart, 2020). Many studies found that this Right-Turn-On-Red (RTOR) change caused more accidents. Unfortunately, these studies concluded that the results were not statistically significant. Only years later, when combining the data, it was found that the changes were significant (Preusser et al., 1982). Statisticians nowadays solve these kind of problems by considering not only significance, but also power. This post aims to explain and demonstrate power from scratch. Specifically, data is simulated to show power and the necessary concepts underlying power. The underlying concepts are the

Show more

2019-12-03

Niceties in the Julia programming language

In general I'm quite amazed by the Julia programming language. This blog post aims to be a demonstration of its niceties. The post targets readers who have programming experience. To aid in the rest of the examples we define a struct and its instantiation in a variable.

struct MyStruct
  a::Number
  b::Number
end

structs = [MyStruct(1, 2), MyStruct(3, 4)]
Functions and methods

For object-oriented programmers the distinction between a function and a method is simple. If it is inside a class it is a method, otherwise it is a function. In Julia we can use function overloading. This means that the types of the input parameters, or signatures, are used to determine what should be called. In Julia these are called methods. For example we can define the following methods for the function f.

Show more

2019-12-01

NixOS configuration highlights

I have recently started paying attention to the time spent on fine-tuning my Linux installation. My conclusion is that it is a lot. Compared to Windows, I save lots of time by using package managers to install software. Still, I'm pretty sure that most readers spend (too) much time with the package managers as well. This can be observed from the fact that most people know the basic apt-get commands by heart.

At the time of writing, I'm happily running and tweaking NixOS for a few months. NixOS allows me to define my operating system state in one configuration file configuration.nix. I have been pedantic and will try to avoid configuring my system outside the configuration. This way, multiple computers are in the same state and I can set-up new computers in no time. The only manual steps have been NixOS installation, Firefox configuration, user password configuration and applying a system specific configuration. (The latter involves creating two symlinks.)

Show more

◀ prev

▶ next